Line of inquiry
Inquiring lines›What determines the reliability an…›How do systems improve effectively…›this line of inquiry
How can evaluation criteria remain robust against agent gaming?
A broader line of inquiry — a family of 29 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 29
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- What makes an evaluation criterion non-stationary enough to resist agent optimization?
- Can a progressively stricter evaluator act like a curriculum for improving agents?
- Can co-evolving evaluators alongside actors prevent reward hacking?
- Does fixed evaluation criteria saturate as self-improving agents improve?
- Can a static evaluator become the performance ceiling for an improving actor?
- Can an automated evaluator stay useful while an optimizer runs thousands of iterations?
- Should feedback channels be excluded from the reward path in agent evaluations?
- Can co-evolved critics truly circumvent static evaluator limitations in self-improvement?
- How does controlled utility evolution prevent the evaluator from becoming a new bottleneck?
- What makes evolving the benchmark different from evolving the optimizer itself?
- Does meta-judging improve evaluator quality better than temporal decoupling alone?
- Why does moving the reward target prevent saturation better than finding a better static proxy?
- Does measured performance gain reflect true task improvement or evaluator exploitation?
- Can an agent change reward-path state through actions during evaluation?
- Should we train the evolver or the executor when building self-improving agents?
- Why do static evaluators become a constraint on model improvement over time?
- Why do automated evaluators enable longer evolutionary loops than human feedback?
- What distinguishes genuine task improvement from evaluator exploitation?
- How do frozen executors with editable text state compare to end-to-end fine-tuning?
- How do you prevent stale reward signals when skills evolve during deployment?
- Can moving or evolving objectives prevent misalignment in discovery agents?
- Why does strengthening the judge improve the actor's generation performance?
- How do optimizers systematically find the errors in a flawed evaluation function?
- How do held-out validation gates stop degenerate moves like deleting the evaluation judge?
- Can a proposer agent actively surface a solver's weaknesses to prevent plateau?
- Can an occasionally wrong judge operate safely in an optimizer loop?
- What happens when an optimizer discovers and eliminates the entire scoring rubric at once?
- How do epoch boundaries preserve self-improvement guarantees across objective changes?
- Why does a series of improving scores differ from a single score rise?