Controlled interventions. Auditable labels.
The benchmark is designed around deterministic replay, strict schema validation, and an explicit held-out generalization test.
Generate
Seeded fictional names, dates, identifiers, and procedural avatars.
Render
One original landscape layout with stored field bounding boxes.
Tamper
Font swap, copy/paste splice, or one-character digit edit.
Degrade
Blur, Gaussian noise, and JPEG compression applied after tampering.
Infer
A versioned prompt and common model-adapter contract.
Evaluate
Extraction, detection, type, localization, robustness, and efficiency.
Strict label contract
Held-out design
The schema rejects any digit_edit example in train or validation. The validator independently checks manifests and requires digit edits in eval_unseen. Seen and unseen results are saved in separate artifacts.
Test the causal intervention, not just the class label.
Counterfactual pairs
Thirty pairs share identity, avatar, layout, and degradation seed. Only the tamper intervention changes, isolating verdict sensitivity and extraction spillover.
Hard negatives
Forty genuine cards contain suspicious but legitimate rendering conditions. False-positive rates reveal whether unusual appearance is being mistaken for manipulation.
Selective review
ECE, Brier score, reliability bins, and risk–coverage curves test whether low-confidence cases can be escalated instead of forcing a verdict.
Metrics
Extraction
Exact match, normalized exact match, character error rate, and fuzzy similarity per field.
Tamper + pairs
F1, false-positive rate, pairwise probability change, correct verdict flips, and outside-field stability.
Uncertainty
Expected Calibration Error, Brier score, reliability bins, abstentions, and selective accuracy at review thresholds.