T+TAFE-IDFORENSICS LABPaper ↗

TAFE-ID / 04 / FAILURE LOG

The weak spots
stay visible.

A forensic tool is only useful when its uncertainty is part of the interface. These are the failure modes recorded during the project.

HELD-OUT PROFILE

Some manipulation families remain difficult.

0.154 F1

Inpaint

Inpainting had the weakest recorded category F1 in this small reference-model check. This metric alone does not establish the cause.

0.234 F1

Insert

Insertion also had low reference-model F1. The subset is too small to isolate the causes or establish a reliable category ranking.

0.036 F1

GroupNorm pilot

The GroupNorm baseline reached only 0.0364 F1 at update 1000 on its held-out subset. Strong fixed-crop scores did not translate into useful document localization.

LIMITS

What the demo does not claim.

The public page does not claim a newly trained state-of-the-art model, a complete paper benchmark reproduction, or production readiness. It exposes a verified reference inference path and the evidence needed to reproduce the decision.