Of 28 depression findings drawn from one public health dataset, only 13 survived a correction for running many tests.
Do depression associations in NHANES papers survive correction for multiple comparisons?
This explores whether the links between depression and various health factors reported in papers built on NHANES (the large US National Health and Nutrition Examination Survey) still hold once you account for how many statistical tests were run, and what that says about how this kind of research is being produced.
This explores whether depression findings drawn from the NHANES public health dataset hold up once you adjust for the number of tests being run. The short answer from the collection: fewer than half do. A systematic review of single-factor NHANES papers re-checked 28 depression studies and found that only 13 of the 28 associations survived correction for multiple comparisons Why did single-factor NHANES studies explode after 2021?. The other 15 are the kind of result you would expect to turn up by chance when enough variables are tested against the same outcome.
The reason this matters is the shape of the studies. A 'single-factor' paper takes one exposure, such as a dietary nutrient, a blood marker or a chemical, and tests whether it is linked to depression. NHANES is public and has hundreds of measured variables, so many teams can each pick a different factor and run the same kind of analysis on the same people. Each paper looks clean by itself, but together they act like one huge fishing expedition. At the usual p < 0.05 threshold, about one test in twenty will look 'significant' even when nothing real is there. Corrections for the false discovery rate exist to catch exactly this, and the review found that most papers left them out. Most also skipped multifactorial modeling, which tests several factors together so that one variable doesn't get credit for what another is doing.
Volume is the other half of the story. These papers averaged about four a year until 2021, then rose to 190 in 2024, with 341 found over the decade Why did single-factor NHANES studies explode after 2021?. When the same template can be applied over and over to a public dataset, the field fills with associations faster than anyone checks them. Other notes in the collection explain why that flood is hard to resist. Readers tend to trust a claim more when it comes with more citations, even irrelevant ones Do users trust citations more when there are simply more of them?. Unreviewed work can shape a debate before anyone checks it Can unreviewed preprints shape scientific debate before peer review?. A pile of weak 'X is linked to depression' papers can start to look like strong evidence just because there are so many of them.
There is a design answer to this in an unexpected place. Spark-to-Paper, a system for generating research papers automatically, requires the evidence standard to be written down before any results are seen. It also separates model judgment from deterministic, checkable steps Can separating judgment from verification improve research paper reliability?. That is essentially pre-registration built into the pipeline. It is the safeguard the NHANES papers were missing: deciding in advance which tests count and correcting for all of them.
The collection has limits here. It has one note that directly addresses NHANES, and that note doesn't say which depression associations survived or which factors they involved. If you want to know whether a particular link, such as vitamin D or sleep, is real, you will need the review itself. The collection does make clear that 'statistically significant in NHANES' should be read as 'worth checking', not as 'established'.
Sources 4 notes
A systematic review found 341 NHANES papers over a decade, with volume jumping sharply after 2021—most omitting false discovery correction and multifactorial modeling. Among 28 depression studies, only 13 of 28 associations survived multiple-comparison correction.
Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.
MIT's case demonstrates that an arXiv preprint shaped AI and science discussions extensively despite never undergoing peer review. When the institution later raised reliability concerns, the damage to discourse had already occurred.
Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review
- Search Arena: Analyzing Search-Augmented LLMs
- Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
- Assuring an accurate research record
- AI for Auto-Research: Roadmap & User Guide
- Seeing to Think? How Source Transparency Design Shapes Interactive Information Seeking and Evaluation in Conversational AI
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries