Evaluator Reliability
Test where learned judges fail, disagree, or become overconfident. I’d build targeted evaluation sets that reveal which answers still require a human decision.
I’ve been exploring what happens when AI systems learn from imperfect evaluators.
A small space for the research problems I’d love to help investigate with you.
As models get more capable, we increasingly rely on other models to evaluate them. That creates a difficult question: what happens when the supervision itself is wrong, overconfident, or easy to game?
I want to identify those failures early, measure how they spread through a training loop, and design escalation rules that bring people in exactly where judgment is uncertain.
better feedback.
→ safer systems.
Test where learned judges fail, disagree, or become overconfident. I’d build targeted evaluation sets that reveal which answers still require a human decision.
Turn raw confidence into a useful decision signal. The goal is to route uncertain cases to people without slowing down every reliable prediction.
Stress-test evaluators under distribution shift, manipulation, and hard negatives. Then trace why their judgments change instead of reporting a single aggregate score.
Turn open-ended questions into reproducible baselines, ablations, and measurements. Every experiment should make the next research decision clearer.
Prioritize scarce human verification using entropy, disagreement, or calibration failure.
Test whether a recursively trained policy learns systematic weaknesses in its evaluator.
Study whether model confidence still means anything under adversarial or out-of-distribution inputs.
Measure how errors propagate when model-generated supervision is reused across training rounds.
Give me a question worth testing, a baseline to break, and people I can learn from.
Same curiosity.
Bigger questions.