RESEARCH DEMO — NOT A PRODUCTION KYC SYSTEM
NNirikshaID Bench
Research methodology

Controlled interventions. Auditable labels.

The benchmark is designed around deterministic replay, strict schema validation, and an explicit held-out generalization test.

01

Generate

Seeded fictional names, dates, identifiers, and procedural avatars.

02

Render

One original landscape layout with stored field bounding boxes.

03

Tamper

Font swap, copy/paste splice, or one-character digit edit.

04

Degrade

Blur, Gaussian noise, and JPEG compression applied after tampering.

05

Infer

A versioned prompt and common model-adapter contract.

06

Evaluate

Extraction, detection, type, localization, robustness, and efficiency.

Strict label contract

{ "doc_id": "nid_000501", "fields": { "...": "clean source values" }, "displayed_fields": { "...": "visible values" }, "tamper_type": "digit_edit", "tamper_bbox": [x1, y1, x2, y2], "split": "eval_unseen", "seed": 20261224 }

Held-out design

The schema rejects any digit_edit example in train or validation. The validator independently checks manifests and requires digit edits in eval_unseen. Seen and unseen results are saved in separate artifacts.

Diagnostic v1.1

Test the causal intervention, not just the class label.

Counterfactual pairs

Thirty pairs share identity, avatar, layout, and degradation seed. Only the tamper intervention changes, isolating verdict sensitivity and extraction spillover.

Hard negatives

Forty genuine cards contain suspicious but legitimate rendering conditions. False-positive rates reveal whether unusual appearance is being mistaken for manipulation.

Selective review

ECE, Brier score, reliability bins, and risk–coverage curves test whether low-confidence cases can be escalated instead of forcing a verdict.

Metrics

Extraction

Exact match, normalized exact match, character error rate, and fuzzy similarity per field.

Tamper + pairs

F1, false-positive rate, pairwise probability change, correct verdict flips, and outside-field stability.

Uncertainty

Expected Calibration Error, Brier score, reliability bins, abstentions, and selective accuracy at review thresholds.

Safeguards

No official symbols, seals, signatures, QR codes, layouts, or real identities are used. Cards are visibly marked synthetic and are unsuitable for identification.