Evaluator Reliability
Test where learned judges fail, disagree, or become overconfident. I’d build targeted evaluation sets that reveal which answers still require a human decision.
I explore what happens when AI systems learn from imperfect evaluators.
Research interests spanning oversight, robustness, calibration, interpretability, and reliable evaluation.
As models get more capable, we increasingly rely on other models to evaluate them. That creates a difficult question: what happens when the supervision itself is wrong, overconfident, or easy to game?
I want to identify those failures early, measure how they spread through a training loop, and design escalation rules that bring people in exactly where judgment is uncertain.
better feedback.
→ safer systems.
Test where learned judges fail, disagree, or become overconfident. I’d build targeted evaluation sets that reveal which answers still require a human decision.
Turn raw confidence into a useful decision signal. The goal is to route uncertain cases to people without slowing down every reliable prediction.
Stress-test evaluators under distribution shift, manipulation, and hard negatives. Then trace why their judgments change instead of reporting a single aggregate score.
Turn open-ended questions into reproducible baselines, ablations, and measurements. Every experiment should make the next research decision clearer.
Prioritize scarce human verification using entropy, disagreement, or calibration failure.
Test whether a recursively trained policy learns systematic weaknesses in its evaluator.
Study whether model confidence still means anything under adversarial or out-of-distribution inputs.
Measure how errors propagate when model-generated supervision is reused across training rounds.
Questions worth testing, baselines worth breaking, and evidence that makes safer systems possible.
Same curiosity.
Bigger questions.