SYNTHESIS NOTE
Topics›Domain Specialization›this note

Why did single-factor NHANES studies explode after 2021?

NHANES papers proposing one-predictor associations surged from 4 per year to 190 in 2024, raising questions about what enabled the volume spike and whether design shortcuts became systematically common.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

Suchak et al. report a systematic literature search that found 341 NHANES-derived papers over the past decade, each proposing an association between one predictor and one health condition. The volume is the headline: "an average of 4 single-factor manuscripts ... per year between 2014 and 2021," then "190 in 2024 up to 9 October." The authorship shift is as sharp: 2 of 25 manuscripts from 2014 to 2020 had a primary author affiliated in China, against 292 of 316 from 2021 to 2024. The depression studies are the authors' test of the statistics. Of 28 depression papers, none applied false discovery correction. When the authors applied Benjamini-Yekutieli correction across the 28 associations, "less than half (13) remained statistically significant."

The mechanism the authors give is about how cheap the work has become. NHANES is an "AI-ready dataset" that can be pulled via API straight into R or Python, so the number of hypotheses is "constrained only by computational access." The authors argue that single-factor AI-supported analysis "removes context from research, fails to capture interactions, avoids false discovery correction, and is an approach that can easily be adopted by paper mills." Two further moves multiply output: reversing predictor and outcome, which they call the most extreme case, and selective data use, such as limited date ranges or cohort subsets, which they read as "suggestive of data dredging, and post-hoc hypothesis formation." The statistics are not new. The claim is that an AI-supported pipeline makes a known practice cheap enough to run at scale.

Against the nearest notes, this excerpt documents the failure that Can separating judgment from verification improve research paper reliability? is designed to prevent. Spark-to-Paper keeps checkable operations apart from model judgment and sets required evidence before results are seen. The depression studies skipped exactly the kind of deterministic step, a multiple-comparison correction, that such a pipeline could enforce. This excerpt is an audit of published output; Spark-to-Paper is a builder's design. The two read as complementary. The surge also sharpens How fast did LLM writing adoption actually spread?. That note measures a rise in public-facing text across four domains; this one shows a rise concentrated in one formulaic genre, with defects in the design rather than the sentences. That matters for Can people reliably spot content made by AI?: a reader relying on prose cues would miss these papers, because the authors' signal is structural (single predictors, no correction, shifting windows).

What the excerpt does not establish is the link to AI itself. The authors measure volume, affiliation and analytic design. They do not detect AI use in any of the 341 papers. "AI-assisted productivity" is an inference from the timing of the rise, and the paper-mill connection is framed as a risk and a case study, not a finding. The China-affiliation shift is a bibliometric fact that says nothing about who produced any given paper. The 13-of-28 result is a reanalysis that treats the 28 depression studies as one family of hypotheses, and it depends on that choice. The excerpt also ends before the authors' best-practice recommendations, so those remedies are not assessed here. The defensible reading is narrower than the headline. The rise in single-factor NHANES papers and their missing corrections are documented. That AI tools caused the rise is a hypothesis the excerpt motivates but does not test.

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What gaps exist between benchmark performance and real deployment outcomes? What human oversight must AI research systems have?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 77 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

NHANES single-factor papers averaged four a year to 2021 and hit 190 in 2024 — formulaic studies skipping multifactorial models and false discovery correction