Line of inquiry
Inquiring lines›How do we evaluate and improve AI…›How can we effectively evaluate AI…›this line of inquiry
How does evaluation scope and dimensionality affect what we measure?
A broader line of inquiry — a family of 40 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 40
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can evaluation trajectories and interaction histories replace single-answer scoring?
- Why does evaluating multiple candidates work better than judging one answer?
- How can interactive evaluation avoid replicating fragmentation problems from response-centered benchmark culture?
- Why does enlarging the evaluation unit reintroduce comparability problems?
- Do current math benchmarks measure outcomes or rhetorical plausibility?
- Can the eight-dimension rubric predict which question types need decomposition?
- What role do multi-dimensional quality frameworks play in assessing arguments versus single-metric approaches?
- Why does item discrimination matter more than surface-level question plausibility?
- Can a single competence score capture multiple separable dimensions of capability?
- Can beam search and ranking functions evaluate claims without understanding counterarguments?
- How does evaluation format change what we measure about model reasoning?
- Can contextual design decisions resist formalization into evaluation rubrics?
- How much does forcing single-choice answers damage alignment with complex intent?
- Does modeling single elicited answers capture real human values or measurement artifacts?
- What happens when we use a single response per condition?
- How does score granularity connect to verification as a scaling axis?
- What makes the Brier score mathematically better than log-likelihood here?
- What separates verifiable reasoning from open-ended judgment in scaling requirements?
- How does unidimensionality in assessments affect measurement validity?
- What compute costs separate a panel of judges from a single large judge?
- How does saturation-aware aggregation encourage balanced improvements across multiple rubric dimensions?
- How does separating environment components make evaluation results more reproducible and analyzable?
- What reporting standards would make interactive evaluation scores comparable across benchmarks?
- How can models select the optimal question to ask given multiple uncertainties?
- How should outcomes be scored when comparing applications with different interaction formats?
- What happens when majority voting converges to a single answer?
- What distinguishes evaluative stance-taking from the mechanical conformity shape-holding describes?
- Can similar outputs from different systems prove they work the same way?
- What specific metrics distinguish single-turn versus multi-turn collaboration success?
- Why does sophisticated measurement not validate the underlying scientific inference?
- What determines whether an answer counts as valid in a particular domain?
- Why is the Judging preference constant while other traits vary slightly?
- What makes a process for choosing between values legitimate and fair?
- What privacy-preserving evaluation methods best capture real-world forecasting ability?
- What would a diagnosable evaluation look like compared to a scalar score?
- How do experts select which other experts to trust?
- Why does a series of improving scores differ from a single score rise?
- How does the evaluator become part of the definition of intelligence?
- What makes proof writing and paper writing harder to verify than proof grading?
- How do you handle disagreement between two accounts of the same incident?