When people lean heavily on AI and don't say so, can detectors catch them, or mostly just the careless?
Do LLM detectors catch undisclosed heavy use at the disclosure threshold?
This explores whether AI-text detectors can reliably catch people who used an LLM heavily without saying so, especially near the line where use becomes heavy enough that it should have been disclosed. The corpus looks at this mostly through how AI conferences policed LLM use in papers and peer reviews.
This explores whether AI-text detectors can catch people who relied heavily on an LLM and didn't disclose it, especially near the point where disclosure becomes required. The short answer from the corpus is no, not on their own. The more useful finding is that the institutions using these tools have largely stopped expecting them to. None of the collected material measures detector accuracy at that exact line, and the place where it is closest is the AI conference world.
The clearest number comes from ICML. It planted hidden instructions in submitted PDFs, a kind of watermark that would show up in a review if a reviewer pasted the paper into an LLM. This flagged 795 reviews, about 1%, and led to 497 desk rejections. The chairs themselves say the method mainly catches careless use. Anyone who noticed and removed the planted instruction, or rewrote the output, got through How many peer reviewers secretly used LLMs despite the ban?. In other words, the detector catches people who copy and paste, not heavy users who edit. The heavy, undisclosed, lightly edited use the question asks about is the case it is weakest on.
ICLR 2026 took the opposite approach. It didn't treat detector flags as verdicts. A flag was one input passed to area chairs, who had to find concrete evidence before any sanction How can conferences detect and handle LLM misuse in peer review?. The program chairs judged authors on two things: whether they disclosed their LLM use, and whether the paper made false claims. The detector score itself wasn't grounds for a penalty How do detection tools shape LLM use enforcement at ICLR?. The thing that was actually enforced was fabricated references. A fake citation can be checked, and a style score can't. So the practical answer to "can we catch undisclosed use?" became "catch the damage it leaves behind." Fake citations are risky for a second reason too: LLM judges score answers higher when they include references, real or not Can LLM judges be tricked without accessing their internals?. That means invented sources can mislead both detection and evaluation.
A comparison from AI safety shows why outside detectors are so limited. When researchers can look inside the model doing the work, cheap "difference-of-means" probes read from the model's internal activations and catch misbehavior about as well as a separate LLM monitor, at almost no cost How do cheap vector detectors compare to expensive LLM monitors?. A conference never has that kind of access to the author's model. It sees only the finished text, after any human editing. That gap helps explain why text detectors struggle while internal monitoring works.
There's also a twist on the disclosure side. One study found that LLM raters quietly favored Black or women authors when AI use went undisclosed, and those preferences disappeared once AI use was disclosed. Human raters gave everyone the same penalty for disclosing Do LLM raters show hidden demographic preferences that disclosure erases?. So disclosure isn't a neutral label. It changes how the work gets judged, by both people and models, and that gives authors near the line a real reason to stay quiet. The takeaway: detectors don't solve this. Institutions are moving to policies built on disclosure, human evidence review, and checkable lapses like fabricated citations.
Sources 6 notes
Hidden-instruction watermarks planted in PDFs flagged about 1% of reviews under ICML's no-LLM rule, leading to 497 desk rejections. The chairs acknowledge the method catches mainly careless uses and misses reviewers who removed or rewrote the watermark.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
ICLR 2026 uses LLM detection tools only to triage papers for area chairs, who must find concrete evidence before sanctions are applied. This human gate protects against false positives by shifting costs from paper rejections to reviewer time.
Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.
On DeepSWE, difference-of-means vectors caught 3.1% more hacks in Kimi K3 but 7.9% fewer in GLM 5.2 than LLM monitors at matched false positive rates. The method applies to existing forward passes, making it virtually free compared to running a separate monitor model.
Show all 6 sources
GPT-4o-mini showed pronounced preference for Black authors and Qwen2.5-7B-Instruct favored women authors when AI use was undisclosed, but both preferences vanished under disclosure. Human raters showed uniform disclosure penalties regardless of author demographics.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- A Retrospective on the ICLR 2026 Review Process
- On Violations of LLM Review Policies
- ICLR 2026 Response to LLM-Generated Papers and Reviews
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review