INQUIRING LINE

Why did human reviewers of an AI theme-sorting tool keep adding missed topics far more than removing wrong ones?

Why did reviewers add themes to Consult's output much more than removing them?

This explores why, when human experts checked the UK government's Consult tool (which sorts public consultation responses into themes), they more often added themes the AI had missed than removed themes it had wrongly assigned, and what that lopsided pattern says about how the AI fails.


This explores why reviewers checking Consult's theme assignments added missing themes more often than they deleted wrong ones. First, the limits of the collection. The one note on Consult, Does AI theme-mapping perform as well as human reviewers?, reports overall agreement: the tool scored F1 0.76 against expert reviewers, while two human reviewers scored 0.81 against each other. The summary doesn't break that gap into themes added versus themes removed. So the collection can't confirm the asymmetry or explain it. Any 'why' below is a reasonable reading, not a documented finding.

Here is that reading. An F1 score blends two kinds of mistakes: labelling something that isn't there, and missing something that is. If reviewers mostly added themes, the tool's errors were mainly misses. It was conservative and tagged a response with fewer themes than a careful human would. That is a different failure from inventing themes, and arguably a safer one. A missed theme leaves a gap someone can fill. A made-up theme quietly distorts the picture. The same note offers a reassuring detail: the disagreements rarely changed which themes ranked at the top. Whatever the AI missed sat mostly in the long tail, not among the main messages.

The review setup itself may also push in one direction. When you check someone else's work, cutting a label means arguing that it's wrong. Adding one only means noticing something extra. Research on polished AI output suggests people tend to trust what is already on the page. Does polished AI output trick audiences into trusting it? and Does polished writing actually signal better quality work? show that confident, professional-looking output leads evaluators to accept it at face value. If reviewers gave Consult's existing labels the benefit of the doubt, removals would be undercounted even where they were deserved. In that case the asymmetry would say as much about the reviewers as about the AI.

There's a useful comparison in peer review. In Can LLM feedback help peer reviewers improve their own reviews?, reviewers who changed their reviews after AI feedback mostly added specific detail rather than taking claims back. That fits the same pattern: when people edit a draft, adding is the easy move. Note that this is an analogy from a different setting, not evidence about Consult.

The main lesson: when a human checks an AI's work, whether they mostly add or mostly remove tells you which way the system errs. It also tells you whether the check is catching both kinds of mistake. If you want a solid answer to the 'why', you'd need Consult's full evaluation report, which this collection doesn't have.


Sources 4 notes

Does AI theme-mapping perform as well as human reviewers?

UK government's Consult tool achieved F1 0.76 against expert reviewers, compared to F1 0.81 between two human reviewers. Differences rarely affected which themes ranked top, suggesting AI performance was competitive despite inherent subjectivity in theme assignment.

Does polished AI output trick audiences into trusting it?

Generative AI produces visually sophisticated outputs without underlying judgment, leveraging the historical heuristic that professional-looking work signals expert thinking. This substitution is especially risky for less experienced workers who lack domain knowledge to evaluate substance beyond form.

Does polished writing actually signal better quality work?

Studies show evaluators perceived AI-generated documents as both human-written and better quality than human submissions. This suggests rhetorical polish misleads judgment and should not serve as a quality signal in evaluation.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.