← NINAD NAIKAI SAFETY RESEARCHEXPLORE RESEARCH ↓
01AI SAFETY

Building safer AI.

I explore what happens when AI systems learn from imperfect evaluators.

Research interests spanning oversight, robustness, calibration, interpretability, and reliable evaluation.

✦
EVALUATEAUDITCALIBRATEIMPROVE
BUILDING AI THAT KNOWS WHEN IT’S WRONG.EVALUATE → AUDIT → CALIBRATE → IMPROVE
02WHY THIS INTERESTS ME

The supervision
problem.

As models get more capable, we increasingly rely on other models to evaluate them. That creates a difficult question: what happens when the supervision itself is wrong, overconfident, or easy to game?

I want to identify those failures early, measure how they spread through a training loop, and design escalation rules that bring people in exactly where judgment is uncertain.

better feedback.
→ safer systems.
03WHERE I CAN CONTRIBUTE
⚖

Evaluator Reliability

Test where learned judges fail, disagree, or become overconfident. I’d build targeted evaluation sets that reveal which answers still require a human decision.

▥

Calibration & Uncertainty

Turn raw confidence into a useful decision signal. The goal is to route uncertain cases to people without slowing down every reliable prediction.

◇

Adversarial Robustness

Stress-test evaluators under distribution shift, manipulation, and hard negatives. Then trace why their judgments change instead of reporting a single aggregate score.

⚙

Experimentation

Turn open-ended questions into reproducible baselines, ablations, and measurements. Every experiment should make the next research decision clearer.

04IDEAS I’D LOVE TO EXPLORE

·audit.selector

Prioritize scarce human verification using entropy, disagreement, or calibration failure.

low riskhigh risksend to human

·judge.game

Test whether a recursively trained policy learns systematic weaknesses in its evaluator.

Policy⇄JudgeCan the model learn
to game the judge?

·confidence.shift

Study whether model confidence still means anything under adversarial or out-of-distribution inputs.

confidenceerror rate

·grounding.loop

Measure how errors propagate when model-generated supervision is reused across training rounds.

Generate→Judge→Train
errors accumulate?
05HOW I’D WORK
▣READMap the claim, assumptions, data, and evaluation protocol.→
▷REPRODUCEBuild the smallest baseline that makes the problem measurable.→
♙TESTUse shifts, adversarial cases, and ablations to find weak points.→
▤WRITEExplain the result, limitations, and the most useful next experiment.
06RESEARCH DIRECTION

I want to work on problems
where the answer isn’t obvious.

Questions worth testing, baselines worth breaking, and evidence that makes safer systems possible.

Same curiosity.
Bigger questions.