RESEARCH DEMO — NOT A PRODUCTION KYC SYSTEM
NNirikshaID Bench
Benchmark results

Seen and unseen failures stay separate.

Normalized exact match preserves meaningful characters. Held-out recall is reported independently, and missing localization predictions are not silently scored as successes.

ModelSplitSamplesField NEMTamper F1Type macro F1Median latency
Qwen2.5-VL-3B-Instruct
zero-shot local VLM
eval_seen2098.0%0.0%22.2%183861.5 ms
Qwen2.5-VL-3B-Instruct
zero-shot local VLM
eval_unseen2099.0%0.0%31.0%180123.1 ms
Template-Aware Tesseract OCR + Visual RF
classical OCR + trained image features
eval_seen2093.0%82.4%72.2%1738.6 ms
Template-Aware Tesseract OCR + Visual RF
classical OCR + trained image features
eval_unseen20100.0%82.4%25.6%1703.9 ms
Qwen2.5-VL-3B-Instruct
ModeZero-shot local inference via OllamaQuantizationOllama Q4_K_MDatasetv1.0.0Spliteval_seenSamples20Promptniriksha-v2

Small diagnostic run: interpret these estimates cautiously.

eval seen · field extraction

full name
100%
guardian name
100%
date of birth
100%
identity number
90%
document id
100%
Localization mean IoU

N/A

JSON parse success

100.0%

Held-out digit-edit recall

N/A

Clean versus degraded

Clean field EMClean supportDegraded field EMDegraded supportPerformance drop
95.0%498.8%16-3.8%

Confusion matrix [TN, FP] / [FN, TP]: [[10,0],[10,0]]

Qwen2.5-VL-3B-Instruct
ModeZero-shot local inference via OllamaQuantizationOllama Q4_K_MDatasetv1.0.0Spliteval_unseenSamples20Promptniriksha-v2

Small diagnostic run: interpret these estimates cautiously.

eval unseen · field extraction

full name
100%
guardian name
100%
date of birth
100%
identity number
95%
document id
100%
Localization mean IoU

N/A

JSON parse success

100.0%

Held-out digit-edit recall

0.0% — 10 samples

Clean versus degraded

Clean field EMClean supportDegraded field EMDegraded supportPerformance drop
97.8%9100.0%11-2.2%

Confusion matrix [TN, FP] / [FN, TP]: [[9,1],[10,0]]

Template-Aware Tesseract OCR + Visual RF
ModeFixed-region OCR; RF fit on train pixels/labelsQuantizationnot applicableDatasetv1.0.0Spliteval_seenSamples20Promptocr-fixed-regions-v1

Small diagnostic run: interpret these estimates cautiously.

eval seen · field extraction

full name
95%
guardian name
85%
date of birth
95%
identity number
90%
document id
100%
Localization mean IoU

N/A

JSON parse success

100.0%

Held-out digit-edit recall

N/A

Clean versus degraded

Clean field EMClean supportDegraded field EMDegraded supportPerformance drop
95.0%492.5%162.5%

Confusion matrix [TN, FP] / [FN, TP]: [[10,0],[3,7]]

Template-Aware Tesseract OCR + Visual RF
ModeFixed-region OCR; RF fit on train pixels/labelsQuantizationnot applicableDatasetv1.0.0Spliteval_unseenSamples20Promptocr-fixed-regions-v1

Small diagnostic run: interpret these estimates cautiously.

eval unseen · field extraction

full name
100%
guardian name
100%
date of birth
100%
identity number
100%
document id
100%
Localization mean IoU

N/A

JSON parse success

100.0%

Held-out digit-edit recall

0.0% — 10 samples

Clean versus degraded

Clean field EMClean supportDegraded field EMDegraded supportPerformance drop
100.0%9100.0%110.0%

Confusion matrix [TN, FP] / [FN, TP]: [[10,0],[3,7]]

The fixed template makes region localization and OCR comparatively easy. Tamper-type generalization remains substantially harder, especially for the held-out digit-edit intervention.

Research diagnostics · v1.1.0

Odd-looking does not necessarily mean tampered.

Hard negatives and causal pairs expose behavior hidden by aggregate F1. All values below are measured from image-only OCR/RF predictions.

42.5%hard-negative false positives · 40 samples
46.7%correct causal verdict flips · 30 pairs
0.053expected calibration error · 120 decisions

F1 under hard negatives

Before
82.4%
After
41.2%

Pairwise intervention

Mean tamper-probability change: 0.290

Outside-field extraction stability: 100.0%

Uncertainty

Brier score: 0.205

10.0% review → 70.4% selective accuracy; 4 errors avoided

20.0% review → 72.9% selective accuracy; 10 errors avoided

30.0% review → 73.8% selective accuracy; 14 errors avoided

Hard-negative false-positive rate by condition

alignment shift · n=5
60%
kerning variation · n=5
100%
low contrast · n=5
20%
printer noise · n=5
20%
scan shadow · n=5
20%
slight font variation · n=5
100%
text compression · n=5
0%
uneven jpeg blocks · n=5
20%

Reliability diagram data

0.50.6 · n=33
57.6%
0.60.7 · n=34
70.6%
0.70.8 · n=26
76.9%
0.80.9 · n=14
64.3%
0.91.0 · n=13
92.3%
Failure analysis

Where image-only systems break

Failure example hardneg_0000

hard-negative false positivenone

hardneg_0000

Legitimate slight font variation was confused with manipulation.

Prediction versus ground truth
{
  "truth": "none",
  "prediction": "font_swap"
}
Failure example hardneg_0002

hard-negative false positivenone

hardneg_0002

Legitimate alignment shift was confused with manipulation.

Prediction versus ground truth
{
  "truth": "none",
  "prediction": "none"
}
Failure example hardneg_0007

hard-negative false positivenone

hardneg_0007

Legitimate kerning variation was confused with manipulation.

Prediction versus ground truth
{
  "truth": "none",
  "prediction": "font_swap"
}
Failure example hardneg_0008

hard-negative false positivenone

hardneg_0008

Legitimate slight font variation was confused with manipulation.

Prediction versus ground truth
{
  "truth": "none",
  "prediction": "font_swap"
}
Failure example hardneg_0010

hard-negative false positivenone

hardneg_0010

Legitimate alignment shift was confused with manipulation.

Prediction versus ground truth
{
  "truth": "none",
  "prediction": "none"
}
Failure example hardneg_0015

hard-negative false positivenone

hardneg_0015

Legitimate kerning variation was confused with manipulation.

Prediction versus ground truth
{
  "truth": "none",
  "prediction": "font_swap"
}
Failure example hardneg_0016

hard-negative false positivenone

hardneg_0016

Legitimate slight font variation was confused with manipulation.

Prediction versus ground truth
{
  "truth": "none",
  "prediction": "font_swap"
}
Failure example hardneg_0017

hard-negative false positivenone

hardneg_0017

Legitimate uneven jpeg blocks was confused with manipulation.

Prediction versus ground truth
{
  "truth": "none",
  "prediction": "none"
}
Failure example hardneg_0018

hard-negative false positivenone

hardneg_0018

Legitimate alignment shift was confused with manipulation.

Prediction versus ground truth
{
  "truth": "none",
  "prediction": "none"
}
Failure example hardneg_0023

hard-negative false positivenone

hardneg_0023

Legitimate kerning variation was confused with manipulation.

Prediction versus ground truth
{
  "truth": "none",
  "prediction": "font_swap"
}
Failure example hardneg_0024

hard-negative false positivenone

hardneg_0024

Legitimate slight font variation was confused with manipulation.

Prediction versus ground truth
{
  "truth": "none",
  "prediction": "font_swap"
}
Failure example hardneg_0027

hard-negative false positivenone

hardneg_0027

Legitimate scan shadow was confused with manipulation.

Prediction versus ground truth
{
  "truth": "none",
  "prediction": "none"
}