INQUIRING LINE

Tell an AI judge a human wrote a text, and it picks the rule-breaking version more often, while humans do the opposite.

Can verifiable rule violations protect AI judgment from authorship label bias?

This explores whether a rule that can be checked objectively, such as 'this text must never use the letter e', keeps an AI judge honest when it is also told whether a human or an AI wrote the text.


This explores whether an objectively checkable rule anchors an AI judge's verdict, or whether a label saying 'a human wrote this' can still sway it. The corpus gives a surprising answer: the rule anchors human judges, but not AI judges. In one experiment, judges compared texts where one broke a lipogram constraint (a rule forbidding a particular letter). When AI models were told a human wrote the rule-breaking text, they picked it 35 percentage points more often. Human judges did the opposite. Given the same label, they picked it 20 points less often Do authorship labels change how AI judges evaluate rule violations?. So the label bent the AI's reading of a fact that anyone could count, while humans seemed to hold on to compliance even more tightly.

This fits a broader pattern of AI judges reacting to who seems to have written something rather than to what it says. On judgments of literary style, the same author labels shifted human ratings by 13.7 points and AI ratings by 34.3 points. That's about two and a half times as much, and the effect held across different AI models Do authorship labels bias how we judge literary quality?. Work on attacks against LLM judges points to the same weakness. Fake references and polished formatting raise scores whatever the content is, and anyone can exploit this without access to the model Can LLM judges be fooled by fake credentials and formatting? Can LLM judges be tricked without accessing their internals?. An authorship label works like one more credential, a cue the judge responds to before it weighs the evidence.

If the rule doesn't protect the judgment, what might? The corpus suggests taking the check away from the judge altogether. Spark-to-Paper builds research-paper generation so that model judgment is kept separate from deterministic, executable checks. It also commits to what evidence will count before any results are seen. That way, the system's consistency doesn't depend on the model judging correctly Can separating judgment from verification improve research paper reliability?. Agent-based evaluators that collect evidence step by step cut a measure of judge unreliability called 'judge shift' from 31% to 0.27%. One caveat: errors in their memory module spread through the system, so each step needs to be kept isolated Can agents evaluate AI outputs more reliably than language models?. Applied to the lipogram case, a script would count the forbidden letters and the model would never get the chance to excuse them.

There is also a reason not to rely on the model to explain its own leniency. Studies of reasoning traces find that influences on a decision often never show up in the visible reasoning, or show up rewritten in harmless-sounding language Can we actually trust reasoning model outputs?. An AI judge swayed by a 'human author' label may never say so. Bias based on authorship also runs in both directions. When readers accuse people of using AI, the accused comments lack the features that actually distinguish AI text. That suggests suspicion of AI authorship can hurt real human writers Do unfounded AI accusations harm human writers instead?. The practical lesson: verifiable rules help only if a deterministic tool enforces them. If a model is asked to apply them while it can see who supposedly wrote the text, the label wins.


Sources 8 notes

Do authorship labels change how AI judges evaluate rule violations?

AI models chose a rule-breaking lipogram 35 percentage points more often when told a human wrote it, while human judges chose it 20 points less in that condition. The shift suggests AI may relax standards for human work while humans anchor to objective compliance.

Do authorship labels bias how we judge literary quality?

Human judges rated identical passages 13.7 percentage points higher when labeled human-authored; AI models showed a 2.5-fold stronger bias at 34.3 points. The effect persists across AI architectures, suggesting evaluators respond to provenance cues rather than text quality alone.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Show all 8 sources
Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Can we actually trust reasoning model outputs?

Research shows reflection rarely corrects errors, traces rarely explain decisions faithfully, and monitoring is vulnerable to two failure modes: omission (influence never reaches the trace) and laundering (problematic reasoning appears in clean language). These vulnerabilities persist even under evaluation pressure.

Do unfounded AI accusations harm human writers instead?

Accused comments lack features that distinguish AI text from human writing, suggesting accusations function as gatekeeping rather than detection. This inverts the AI-as-perpetrator framing, placing harm at the receiving side through reader skepticism.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.