M4 · Frozen test split
Semantic + Texture
Accuracy68.18%
Macro-F10.679
Parameters28,355,317
Failures14/44
Learned fusion weights: Semantic 0.503779 · Texture 0.496221. Not causal feature-importance estimates.
61711214012
Rows: true · Columns: predicted


Side-by-side
| Metric | M3 · Semantic | M4 · Semantic + Texture | M5 · Semantic + Edge |
|---|---|---|---|
| Accuracy | 0.591 | 0.682 | 0.636 |
| Macro-F1 | 0.590 | 0.679 | 0.640 |
| BONA_FIDE F1 | 0.414 | 0.480 | 0.480 |
| PRINT F1 | 0.690 | 0.889 | 0.846 |
| SCREEN F1 | 0.667 | 0.667 | 0.595 |
| Failures | 18 | 14 | 16 |
ConvNeXt + Texture produced the highest observed score on this frozen test split. With only 44 test samples and one primary seed, the difference is descriptive rather than evidence of statistical superiority.