A website says fewer humans are visiting — but is that true, or just better bot-blocking making it look that way?
How much of the measured decline is a bot detection artifact?
This explores whether a reported drop (most likely in human activity or traffic on the web) is a real change in behavior or partly an artifact of how bots get detected and filtered out of the count.
This explores whether a measured decline reflects real change or is partly produced by how bot traffic gets identified and removed before counting. The short answer: these retrievals don't speak to bot detection at all, so the collection can't put a number on how much of any decline is artifact. What it does have is a set of cases where a measuring instrument quietly shaped the result. Those cases show what to check before you trust a decline figure.
The clearest parallel is hallucination detection. Progress there looked real until researchers swapped the scoring metric and found that up to about 46% of the apparent gains vanished. Simple length heuristics turned out to match sophisticated methods, so the metric had been tracking something other than what everyone assumed (Is hallucination detection progress real or just metric artifacts?). Bot-filtered traffic data can fail the same way. If the classifier gets stricter, or starts flagging some kinds of human sessions as automated, the 'human' count falls even if human behavior hasn't changed. Reward hacking research makes the same point in general form: until your detector is reliable, you can't tell whether a change in the numbers reflects the world or the instrument. Measurement has to be fixed first (Can we measure reward hacking reliably enough to act on it?).
Detectors also degrade under pressure from the things they detect. One study shows attackers using scanner feedback to tune each piece of their activity until it looks harmless, reaching 96% evasion across six scanners (Can attackers evade skill scanners by refining individual skills?). Another shows that repeated quiet probing can separate decoys from real objects only when their response patterns actually differ (Can repeated quiet probes separate decoys from genuine objects?). Applied to bots, this cuts both ways. If bots get better at passing as humans, measured human activity is inflated and a real decline is hidden. If filters get more aggressive, a decline can appear that isn't there. Without knowing how the detector changed over the measurement period, you can't tell which direction the error runs.
Two more notes are useful for reading any decline claim. Shlegeris's critique of an OpenAI measurement shows that a flat or modest aggregate number can set an upper limit on an effect without ruling out a smaller, targeted one hiding underneath (Can OpenAI's measurements rule out subtle goal suppression?). Nielsen's critique of AI 'confiscation' studies is a reminder that a study can measure its outcome correctly and still measure the wrong outcome, one that doesn't match real conditions (Does removing AI tools actually measure real skill loss?). For any decline you're evaluating, ask three things. Did the bot filter change during the window? Which way would its errors push the count? Does the metric track the behavior you actually care about? If you want direct evidence on bot detection and traffic data, this collection doesn't have it yet.
Sources 6 notes
ROUGE-based evaluation inflates detection capability by up to 45.9 percent compared to human-aligned metrics. Simple length heuristics rival sophisticated methods like Semantic Entropy, suggesting much reported progress measures length variation rather than factual accuracy.
The paper argues that current detection methods are too unreliable to support readiness judgments. Mitigation strategies cannot be properly evaluated until measurement instruments are fixed, making measurement a prerequisite for both safety assessment and deployment decisions.
ColluSkill combines chain planning with scanner-feedback refinement to reach 96% average attack success. The approach works because scanners score skills individually, allowing feedback to reduce suspicion per skill while chain-level semantics remain intact.
In idealized settings with independent responses, enough quiet probes let a classifier separate decoys from genuine objects with vanishing error if their response distributions differ and are known or learnable from feedback.
Shlegeris argues OpenAI's measurements establish an upper bound on CoT-access harms but do not exclude small, targeted suppression of misaligned-goal mentions. A model could learn incidentally to hide specific goals while aggregate monitorability scores remain flat.
Show all 6 sources
Nielsen argues that removing AI tools to test skill retention replicates a scenario outside the research lab, making these studies measure the wrong outcome. He proposes instead studying how higher-level skills develop when AI remains available permanently.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks
- Recent Frontier Models Are Reward Hacking
- Evaluation Awareness Is Not One Capability: Evidence from Open Language Models
- Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
- The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
- LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
- ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners
- The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs