Inside a real screening process, nobody can reliably tell which submissions came from AI, and the people submitting adapt to the filter.
How do live screening workflows differ from controlled experiments with labeled AI output?
This explores what changes when AI-generated material is judged inside a real screening process, such as job application filters or conference peer review, where nobody is told what is AI, compared with a controlled study where the AI output is known and labeled.
This explores what changes when AI-generated material is judged inside a real screening process, where the screener isn't told what's AI, compared with a setup where the AI output is known and labeled. The short version from this corpus: in live screening, the main difficulty is not judging quality. It's that nobody can reliably tell which submissions came from AI, and the people submitting can adapt to the filter.
Start with the detection problem. A review of 30 studies found that people's ability to spot AI-generated text, images and voice clusters around chance, and it hasn't kept up as AI output became more realistic Can people reliably spot content made by AI?. In a labeled experiment, that doesn't matter, because the researcher already knows which items are AI and can measure how judges respond to them. In a live workflow, it decides everything. Screeners are judging blind whether they want to or not. Hiring shows what happens next. Greenhouse's survey found 41% of job seekers using prompt injections (hidden instructions meant to manipulate the AI filters) and 34% of recruiters spending half their week filtering spam. Applicants game the filters and employers tighten them, and neither side ever sees a labeled version of the other Are job applicants and employers locked in an escalating AI arms race?. That note also shows the cost of studying live settings: the data supports each step of this escalation loop, but it can't establish which side drives the other or how the trend moves over time. A controlled design would be able to answer that.
The most useful cases here sit between the two designs. Sakana's AI Scientist-v2 sent three fully AI-generated papers into a real ICLR 2025 workshop review. The reviewers didn't know which papers were AI-written, but the organizers did, and they had agreed in advance to withdraw any paper that was accepted. One paper averaged 6.33 and was withdrawn Can AI-generated papers pass peer review undetected? Can AI systems generate research papers that pass peer review?. This is the live-screening version of a labeled experiment: unlabeled for the judges, labeled for the experimenters. It also shows what blind review misses. The authors later found a citation error, and they concluded that none of the three papers met main-conference standards. The paper got through a real screen without being good enough. The earlier AI Scientist ran its review loop internally instead, with an ensemble of AI reviewers scoring against NeurIPS guidelines Can one AI system complete a full research cycle end-to-end?. That's a controlled setting where the evaluators are built in and everyone knows what they're looking at.
Two more ideas from the corpus help explain why live and controlled results diverge. AI output varies with the prompt, the context and the reader, so a single labeled sample in an experiment may not represent what reaches a live screener Why does AI output change with every prompt and context?. And how much you can trust the evaluation depends on how it's done. An agent that gathers evidence while judging drifted about 100 times less than a one-pass LLM judge on complex tasks Can agents evaluate AI outputs more reliably than language models?. That points to one possible fix for live screening: treat judgment as a process of checking evidence, not a single verdict.
This set of retrievals has a clear gap. It contains no controlled experiment in which judges are told "this is AI" and researchers compare their reactions with an unlabeled condition, so it can't say how labeling itself changes trust, scrutiny or scoring. The takeaway the corpus does support is unexpected: the cleanest evidence about AI in live screening comes from hybrid studies where the organizers hold the label and the judges don't.
Sources 7 notes
A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.
Greenhouse's survey found 49% of job seekers submit more applications than before, 41% use AI prompt injections to bypass filters, while 91% of recruiters spot deception and 34% spend half their week filtering spam. The data supports each leg of the loop but does not establish causal direction or measure the trend over time.
Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.
AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.
The AI Scientist performed ideation, coding, experiments, writing, and self-review autonomously, producing a manuscript that passed the first round at a machine learning workshop with 70% acceptance rate. Five ensemble reviewers and an area-chair model judged the output against NeurIPS guidelines.
Show all 7 sources
AI outputs exhibit essential mutability—they vary with sampling, prompt wording, and audience interpretation. This is not a defect but a defining feature of tokens as media, making them fundamentally different from fixed commodities and resistant to traditional quality assurance.
Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- The AI Scientist Generates its First Peer-Reviewed Scientific Publication
- AI for Auto-Research: Roadmap & User Guide
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
- Stop Automating Peer Review Without Rigorous Evaluation
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Predicting Empirical AI Research Outcomes with Language Models