300 / 300
Successful crops
Dataset + evaluation
The reduced benchmark was designed around provenance, lineage, leakage prevention, and a strict separation between validation decisions and test reporting.
The full acquisition was impractical inside the 24-hour storage and network envelope. Protocol A-Reduced is a controlled 300-sample subset, not the full DLC-2021 benchmark or an equivalent paper reproduction.
DLC-2021 licensing and acquisition constraints were gated before processing. A source-confounded mixed-dataset four-class design was rejected.
Deterministic sampling retained one selected frame per capture and balanced 100 BONA_FIDE, 100 PRINT, and 100 SCREEN samples.
All captures and derivatives share base_document_id lineage. Group-aware splitting produced 214 train, 42 validation, and 44 test samples.
No exact-duplicate clusters, pHash-distance-6 near-duplicate clusters, or base-document split violations crossed splits.
RetinaFace → CLAHE + RetinaFace → Haar fallback → context expansion → 384×384 crop. All 300 crops succeeded.
Train-only class weights, no stochastic augmentation, AdamW, mixed precision on Tesla T4, physical batch 8 with accumulation 4.
Validation Macro-F1 controlled scheduling, early stopping, and checkpoint choice. Test was isolated until selection finished.
Successful crops
RetinaFace first-pass overall; one SCREEN used Haar fallback.
Observed peak RAM during portrait extraction.
Padding was uneven: BONA_FIDE 39%, PRINT 32%, SCREEN 61%. It is documented as a possible shortcut risk, not ignored.