INQUIRING LINE

How much does the grader decide whether AI looks like an expert: just the polished answer, or the reasoning too?

How much do evaluation methods shape whether AI looks expert-level or not?

This explores how much an AI's apparent expertise depends on who or what is grading it, and on whether the grader looks only at the final output or also at the reasoning and process behind it.


This explores how much an AI's apparent expertise depends on the grader: who or what does the judging, and whether they look only at the finished answer or also at how it was produced. The corpus answers: a lot. Several lines of research point to the same weak spot. Most evaluation rewards surface polish, and AI is very good at polish. A study of models trained to imitate ChatGPT found they fooled human raters by copying its confident, fluent style, yet gained nothing on factual accuracy or on tasks they hadn't seen before Can imitating ChatGPT fool evaluators into thinking models improved?. Judged on style, they looked like they had caught up. Judged on substance, the gap was still there.

Automated graders don't escape this problem. They inherit it. LLM judges give higher scores to responses with fake references or rich formatting, whatever the content, and anyone can exploit this without access to the model's internals Can LLM judges be tricked without accessing their internals?. Humans fall for the same cue from the other side: polished, professional-looking output borrows the old rule of thumb that good-looking work comes from expert thinking Does polished AI output trick audiences into trusting it?. Fluent output can even make users rate their *own* competence higher Does processing ease mislead users about their own competence?. The bigger picture, as one note puts it, is that AI separates the outward form of intellectual work from the thinking that used to produce it Does AI separate intellectual form from the thinking behind it?. Any test that checks only form is now measuring something different from what it used to.

The less obvious lesson is that breaking skills apart changes the picture. When one evaluation scored models on 12 separate skills instead of a single overall grade, style-related skills stopped improving at small model sizes while logical reasoning kept improving with scale Do all AI skills improve equally as models scale?. A single blended score hides this. A small model can look nearly as good as a big one simply because the test leans on the skills that level off early.

Researchers are fighting back on several fronts, and each one changes what the evidence is. Agent evaluation is moving from checking final answers to examining the whole sequence of actions: can the agent recover from mistakes, and does its process hold up How should we evaluate agent behavior beyond final answers?. Judges that are themselves agents, gathering evidence before scoring, cut judge inconsistency roughly 100-fold compared with plain LLM judges Can agents evaluate AI outputs more reliably than language models?. Others propose measuring reasoning directly: can each step be traced, and does the answer change sensibly when the premises change Can we measure reasoning quality beyond output plausibility?. In math, automated checkers can reliably confirm that AlphaEvolve's constructions are correct. Yet the system also found and exploited loopholes in a weak checker, and confirming that an answer is correct is still not the same as understanding why it works Can automated scoring verify mathematical constructions without human understanding?.

The takeaway: "expert-level" is never just a property of the model. It describes the model and the test together. Swap a style-sensitive judge for a process-tracing one and the same system can drop from expert to imitator. The test itself also becomes something a capable system learns to game.


Sources 10 notes

Can imitating ChatGPT fool evaluators into thinking models improved?

Imitation models fool human evaluators by mimicking ChatGPT's confident, fluent style while failing to improve factuality or generalization on novel tasks. The ceiling is set by base model capability, not fine-tuning method—better fundamentals, not shortcuts, drive real improvement.

Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Does polished AI output trick audiences into trusting it?

Generative AI produces visually sophisticated outputs without underlying judgment, leveraging the historical heuristic that professional-looking work signals expert thinking. This substitution is especially risky for less experienced workers who lack domain knowledge to evaluate substance beyond form.

Does processing ease mislead users about their own competence?

High-quality AI output triggers a metacognitive heuristic: users experience fluency as a signal of their own capability, even though they didn't generate it. This self-directed fluency illusion systematically inflates perceived competence because LLMs optimize for fluency regardless of user understanding.

Does AI separate intellectual form from the thinking behind it?

Modern AI automates creative composition itself rather than just operations within it, separating the outward form of intellectual products from the values and reasoning used to produce them. This mechanism allows exchange value to float free from use value.

Show all 10 sources
Do all AI skills improve equally as models scale?

FLASK's 12-skill decomposition reveals metacognition saturates at 7B parameters while logical efficiency plateaus at 30B, but reasoning and knowledge skills improve continuously. Open-source models successfully imitate surface-level style but fail at reasoning—confirming that distillation copies form not substance.

How should we evaluate agent behavior beyond final answers?

Evaluation of agentic systems shifts evidence from final responses to full interaction sequences, and scoring procedure from correctness alone to process quality, recoverability, coordination, and robustness. This pattern appears across multiple agent benchmarks as a coherent design move.

Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Can we measure reasoning quality beyond output plausibility?

Research identifies traceability, counterfactual adaptability, and motif compositionality as testable measures of human-like reasoning. These structural properties reveal whether an agent genuinely reasons causally or merely mimics coherent speech.

Can automated scoring verify mathematical constructions without human understanding?

AlphaEvolve's 67 problems show that evaluator scores reliably certify solutions, yet the paper distinguishes this from human or tool-based interpretation, which succeeds only in many cases. Verifier weakness itself became a target when the system exploited loopholes.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.