INQUIRING LINE

What exactly makes a human expert catch mistakes about their field that even a smart AI judge misses?

What makes expert domain knowledge valuable for AI model evaluation?

This explores why AI evaluation needs people (or encoded rules) that really know a field, instead of generic checks or another general-purpose model acting as judge. The corpus has little that targets evaluation head-on, but it says a lot about what domain experts catch that other checks miss.


This explores why AI evaluation needs people (or encoded rules) that really know a field, instead of generic checks or another general-purpose model acting as judge. One caveat first: the collection is mostly about putting expert knowledge *into* models, not about using it to *grade* them. Read sideways, though, it gives a clear answer. The failures that matter most in specialized AI are the ones that look fine to anyone who isn't an expert.

Start with the kind of error at stake. Research on domain specialization finds that models without enough specialist knowledge don't fail loudly in high-stakes settings. They produce confident, plausible-sounding errors How do you build domain expertise into general AI models?. A related finding goes further: domain training often shows visible gains on benchmarks while quietly eroding things you can't see in a score, such as whether the model's stated reasoning matches what it actually did, and whether its skills carry over to new formats How do domain training techniques actually reshape model behavior?. A generic evaluator sees the gain and misses the damage. Spotting a fluent wrong answer, or a correct answer reached for the wrong reasons, takes someone who knows what right looks like in that field.

That explains why plain LLM-as-a-judge setups struggle. One study found that general model judges drifted in their verdicts 31% of the time on complex tasks. An agent-based judge that actively gathered evidence before deciding cut that drift to 0.27% Can agents evaluate AI outputs more reliably than language models?. The lesson carries to expertise: good evaluation means checking claims against something solid, not judging how convincing they sound. A related limit applies too. No amount of clever prompting can give a model knowledge it never learned, only rearrange what it already has Can prompt optimization teach models knowledge they lack?. So a judge model without the domain knowledge can't be prompted into a reliable domain grader.

The more surprising thread is that expert judgment doesn't have to stay locked inside experts' heads. An industrial case study wrote specialists' rules and design principles directly into an AI agent's working setup, and non-experts using it produced work rated at expert level, with no specialist reviewing each output Can codified expertise let non-experts match specialist output?. Training research shows the same pattern. Rewards that grade the *quality of reasoning*, not just whether the final answer is right, help models absorb domain knowledge better than plain fine-tuning does Can reinforcement learning embed domain knowledge more effectively than supervised fine-tuning?. Rubric-based rewards that score intermediate reasoning steps, and are applied only when the final answer is correct, make it hard for a model to game the grader Can search agent behavior yield reliable process rewards for reasoning?. In both cases, expert knowledge acts as a standard for what good *process* looks like, not just a list of right answers.

The deeper reason this matters: systems that learn only from data, with no explicit rules, pick up statistical shortcuts that nothing corrects, and they break outside familiar territory Does refusing explicit knowledge harm AI system performance?. Expert evaluation is often the only place those explicit standards enter the loop. So expert knowledge does more in evaluation than serve as an answer key. It gives you a way to tell "sounds right" from "is right," and the most useful finding here is that this ability can be written down, turned into a rubric, and handed to an automated grader.


Sources 8 notes

How do you build domain expertise into general AI models?

Research shows that over-specialized models fail catastrophically outside their domain, while under-specialized ones produce confident-sounding errors in high-stakes settings. The tension is structural, not solvable through technique alone.

How do domain training techniques actually reshape model behavior?

Research shows every adaptation method—from parameter-efficient tuning to knowledge graph curricula—has optimal conditions tied to specific domains. The key finding: visible benefits like performance gains often come with hidden degradation in reasoning faithfulness, capability transfer, and format flexibility.

Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Can prompt optimization teach models knowledge they lack?

Prompting works entirely within a model's pre-existing training distribution and cannot supply domain knowledge absent from training data. This creates a hard ceiling: no prompt strategy can compensate for missing foundational knowledge, only reorganize what already exists.

Can codified expertise let non-experts match specialist output?

An industrial case study embedding domain rules and design principles into an LLM agent's scaffolding achieved 206% output-quality improvement and expert-level ratings from non-experts, bypassing the need for specialist oversight. The capability gain came from externalizing tacit expertise into structured harness components, not from model scale.

Show all 8 sources
Can reinforcement learning embed domain knowledge more effectively than supervised fine-tuning?

RLAG rewards both answer accuracy and explanation rationality by cycling between augmented and unaugmented generation, progressively internalizing coherent knowledge structures. This outperforms SFT because it prioritizes reasoning quality over token-level correctness.

Can search agent behavior yield reliable process rewards for reasoning?

LongTraceRL mines entity-level reasoning signals from what search agents read but don't cite—the hardest distractors—and applies rubric rewards only to correct answers, structurally blocking reward fabrication while capturing intermediate reasoning quality.

Does refusing explicit knowledge harm AI system performance?

AI systems that learn exclusively from data produce uninterpretable representations, inherit statistical biases uncorrected by normative rules, and fail to generalize beyond training distributions. Structured knowledge injection at minimal corpus cost substantially improves performance.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.