Normalization and loss
A baseline with GroupNorm and a balanced loss learned four fixed crops: F1 ranged from 0.919 to 0.993 at update 200. Save/reload produced identical masks. This checks memorization and serialization only.
TAFE-ID / 03 / TRAINING
We ran bounded custom-model experiments and separately verified a released reference checkpoint. The reference scores are not the result of our fine-tuning.
OUR EXPERIMENTS
A baseline with GroupNorm and a balanced loss learned four fixed crops: F1 ranged from 0.919 to 0.993 at update 200. Save/reload produced identical masks. This checks memorization and serialization only.
The separate GroupNorm pilot reached F1 0.0309 at update 500 and 0.0364 at update 1000 on 32 held-out documents. These low scores did not justify presenting it as a usable detector.
The later TAFE frequency experiment passed a finite-gradient forward/backward check. Its final four-crop F1 ranged from about 0.754 to 0.948. These are diagnostic crop scores, not full-document performance.
REFERENCE CHECKPOINT
The ASCFormer checkpoint was released by the RTM authors. We verified strict loading, CUDA operators and inference in an isolated Python 3.8 environment. The recorded 32-document reference check reached F1 0.623. No claim is made that our custom training improved that checkpoint.
Model research is paused. The current work is documenting the evidence and making the project inspectable.