T+TAFE-IDFORENSICS LABPaper ↗

TAFE-ID / 02 / METHODOLOGY

Data first.
Then the model.

The run was designed around the RTM official split, reproducible inputs and visible stop conditions. Every claim on this site is tied to a recorded artifact.

DATA AUDIT

What was actually prepared.

9,000

RTM images

Images and aligned segmentation masks were restored from the public RealTextManipulation archive.

5,803 / 3,197

Official split

The archive’s train and test lists were preserved. A fixed 32-document verification subset was used for project verification.

DCT + Q-TABLE

Frequency inputs

JPEG coefficient ranges, non-zero fractions and quantization tables were checked before any frequency-aware training.

EVALUATION

Reference result protocol.

The released ASCFormer checkpoint was loaded in the pinned Python 3.8.20, PyTorch 2.0.0, CUDA 11.8, MMCV 2.0.0 and MMEngine 0.7.0 environment. Strict loading, custom CUDA operators and inference self-tests passed. The reported reference metrics come from the fixed verification subset; the custom training branch is reported separately.

0.623 F1 · 0.452 IoUReference inference on 32 RTM verification documents