Theme of inquiry
What architectural and training strategies optimize model efficiency and performance?
A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.
30 specific questions
- Why does evaluating multiple candidates work better than judging one answer?
- Can evaluation trajectories and interaction histories replace single-answer scoring?
- How can interactive evaluation avoid replicating fragmentation problems from response-centered benchmark culture?
- Why does enlarging the evaluation unit reintroduce comparability problems?
- How do ensemble methods reduce bias in automated evaluation?
- Can judges trained on both verifiable and non-verifiable tasks transfer across domains?
- Does ensembling smaller judges reduce bias more effectively than single large judges?
18 specific questions
- Why do explicit ratings fail to capture uncertainty in user preferences?
- How can consistency across measurement conditions identify genuine versus constructed preferences?
- Why does preference measurement validity matter more than aggregation methods?
- What distinguishes genuine user preferences from similar-user preferences in sparse data?
- What consistency tests could distinguish constructed from genuine preferences?
- How do implicit signals like clicks capture preference more reliably than explicit ratings?
- When does low-dimensional preference factorization miss important user variation?
20 specific questions
- Can standard accuracy metrics miss the real constraints on user consumption?
- Why do standard accuracy metrics miss set-level composition constraints in recommendations?
- Why is evaluating synthetic data quality so ambiguous and context-dependent?
- Why do standard accuracy metrics ignore set-level consumption constraints?
- What metrics capture whether recommendations reflect a user's full taste range?
- Why do ranking metrics fail to capture distributional properties of user taste?
- What measurement artifacts emerge when annotators interpret the same question differently?
18 specific questions
- Why does majority voting reward work better than other test-time aggregation methods?
- How does training-time voting differ from inference-time majority voting over samples?
- How does majority voting fail when reasoning samples lack genuine diversity?
- Does majority voting reliably signal correctness without risking reward hacking?
- When does multi-agent voting help versus hurt performance on tasks?
- Can test-time voting improve reasoning beyond the base model's original capabilities?
- Can voting work at every level of task decomposition, not just whole problems?
35 specific questions
- Should AI alignment use normative standards instead of aggregate preferences?
- Does a single LLM judge capture diverse human preferences in alignment training?
- Can communication problems and optimization problems be addressed with the same alignment approaches?
- Can AI-assisted alignment eventually solve fairness at scale?
- Should LLMs align with social roles instead of individual preferences?
- Why do standard preference alignment methods fail at the individual user level?
- Why does RLHF alignment reduce the diversity of viewpoints in AI output?