Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›What mechanisms determine whether…›this line of inquiry
What makes imperfect LLM judges safe for optimization?
A broader line of inquiry — a family of 18 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 18
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- What design choices make it survivable when an LLM judge holds final authority over an optimizer?
- Do mechanical guardrails around judges bound the cost of judge errors?
- How can deterministic checks make wrong judge decisions survivable?
- Can an occasionally wrong judge operate safely in an optimizer loop?
- What makes a judge's calibration at decision boundaries harder to improve?
- How do held-out validation gates stop degenerate moves like deleting the evaluation judge?
- When does an LLM judge's error become terrain that an optimizer can map and exploit?
- What signals could refinement loops exploit in defense verdict systems?
- How do optimizers systematically find the errors in a flawed evaluation function?
- What did deleting the rubric do to the judge's error in practice?
- Can verification loops and decomposition fix judgment failures?
- Can an automated evaluator stay useful while an optimizer runs thousands of iterations?
- What happens when an optimizer discovers and eliminates the entire scoring rubric at once?
- Can an optimizer that sees guardrail verdicts learn to route around them?
- How should unarguable checks order themselves before arguable verification steps?
- How does bounding a judge's authority differ from improving the judge itself?
- Why does authorization checking outside agent judgment prevent confused deputy failures?
- How do held-out gates compare as defenses when the proposer is an LLM?