Paper + feasibility audit
Separated paper-described components from implementation choices before writing model code.
Scope lockedFrom paper audit to public record
The sprint was less about chasing a score than making defensible decisions under dataset, storage, compute, and time constraints.
Separated paper-described components from implementation choices before writing model code.
Scope lockedFound the proposed composite benchmark was source-confounded and treated it as a blocker.
Four-class merge rejectedStreamed only selected frames, preserved lineage, and produced a balanced 300-sample protocol.
214 / 42 / 44Serial decoding, local batch 1, zero workers, no RAM cache, CPU extraction, temporary free T4 training.
~573 MiB extraction peakRetinaFace with CLAHE and Haar fallbacks generated 300 context-preserving 384×384 crops.
300 / 300 cropsConvNeXt-Tiny established a controlled semantic baseline.
0.590 Macro-F1A 64-D pixel-only texture path lifted PRINT F1 from .690 to .889.
0.679 Macro-F1Fixed Sobel gradients added different corrections, but SCREEN F1 fell in this run.
0.640 Macro-F1Compared aggregate metrics and failure transitions without test-set retuning.
5 failures persistedPublished safe code and aggregate evidence while excluding restricted document imagery.
Open research recordA scientifically clean protocol mattered more than immediate training.
Texture and edge changed which samples failed—not simply how many.
Forty-four test samples cannot support broad performance claims.