Explosion of formulaic research articles, including inappropriate study designs and false discoveries, based on the NHANES US national health database

Paper · Source
Domain Specialization in LLMs

Source: Suchak et al., PLOS Biology · 2025-05-08

With the growth of artificial intelligence (AI)-ready datasets such as the National Health and Nutrition Examination Survey (NHANES), new opportunities for data-driven research are being created, but also generating risks of data exploitation by paper mills. In this work, we focus on two areas of potential concern for AI-supported research efforts. First, we describe the production of large numbers of formulaic single-factor analyses, relating single predictors to specific health conditions, where multifactorial approaches would be more appropriate. Employing AI-supported single-factor approaches removes context from research, fails to capture interactions, avoids false discovery correction, and is an approach that can easily be adopted by paper mills. Second, we identify risks of selective data usage, such as analyzing limited date ranges or cohort subsets without clear justification, suggestive of data dredging, and post-hoc hypothesis formation. Using a systematic literature search for single-factor analyses, we identified 341 NHANES-derived research papers published over the past decade, each proposing an association between a predictor and a health condition from the wide range contained within NHANES. We found evidence that research failed to take account of multifactorial relationships, that manuscripts did not account for the risks of false discoveries, and that researchers selectively extracted data from NHANES rather than utilizing the full range of data available. Given the explosion of AI-assisted productivity in published manuscripts (the systematic search strategy used here identified an average of 4 papers per annum from 2014 to 2021, but 190 in 2024–9 October alone), we highlight a set of best practices to address these concerns, aimed at researchers, data controllers, publishers, and peer reviewers, to encourage improved statistical practices and mitigate the risks of paper mills using AI-assisted workflows to introduce low-quality manuscripts to the scientific literature.

The quantity of biological data available to researchers has increased dramatically in recent years, leading to more opportunities for data-driven research. As more information becomes available in artificial intelligence (AI)-ready formats, research—when performed in line with best practices—should become faster and more reproducible. The wide availability of such datasets can, however, introduce new problems, by facilitating end-to-end AI-supported manuscript production on a large scale. This is a practice which may be adopted by paper mills, defined by the United2Act Research Working Group as covert organizations that provide low-quality or fabricated manuscripts to paying clients [1].

Here, we systematically investigate research papers which used data extracted from the National Health and Nutrition Examination Survey (NHANES) [2], a cross-sectional data source originally established to assess the health and nutritional status of adults and children in the United States. NHANES combines interviews, physical examinations and laboratory tests to collect comprehensive data on the prevalence of diseases, risk factors, and health trends. NHANES surveys are conducted on a two-year cycle and aim to recruit around 10,000 participants per survey. While many variables are collected on a continuous basis, others have been included or excluded at different points, as areas of interest to stakeholders have changed, and participants are freshly recruited for each survey. The most recent NHANES survey for which data are available (covering 2021–2023) included over 700 variables. In terms of accessibility, NHANES is an AI-ready dataset, in line with the criteria set out by the NIH Bridge to Artificial Intelligence (Bridge2AI) Standards Working Group [3].

The ability to extract data via an API directly into machine learning environments such as R or Python can transform productivity, with the number of hypotheses that can be tested constrained only by computational access, but this can also carry risks. A focus on single-factor analyses can be especially problematic, given the multifactorial nature of many illnesses, as well as the challenge of differentiating between predictors that are specific to a health condition versus those shared across different disease types [9]. In addition, the ability to generate large numbers of machine-learning models allows for rapid post-hoc investigation of alternative hypotheses, should the main a priori hypothesis not be supported (a form of hypothesizing after the results are known, or HARKing) [10–12]. With ready computational access, it is possible to conduct a broad search for any combination of indicator, health condition, cohort, and time window that yields a low p-value. While data dredging is a well-described phenomenon [13–16], direct-to-AI pipelines can make formulaic research pipelines more productive than has previously been possible. This productivity gain is likely to be particularly attractive to paper mills.

In this work, we conducted a systematic literature search over the last 10 years to retrieve potentially formulaic papers analyzing NHANES data, and analyzed these manuscripts for common themes around statistical approaches, study design, or results that were not translational in nature. We also aimed to identify whether these issues provided a case study of the risks of AI-supported workflows being adopted by paper mills, from workflows to automate data dredging and machine learning, to manuscript preparation using generative AI.

The 341 identified reports were published across a number of different journals (147 journals in total); all articles were successfully retrieved. The top 10 journals accounted for 43% of the articles. This is also shown in a tree map format in Fig 2 below, with the full data in S1 Data Table A. The average impact factor of the journals publishing these papers was 3.6. Three journal families accounted for over half of the manuscripts identified in this review: Frontiers Media SA (22%), BioMed Central Ltd (18%) and Springer (13%).

In terms of trends over time, an average of 4 single-factor manuscripts identified by the search strategy were published per year between 2014 and 2021, increasing rapidly from 2022, with 190 in 2024 up to 9 October. There has also been a general increase in health data-driven research (Fig 3B shows this trend for the search term ‘biobank’ as a simple example), but this wider growth does not explain the scale of the increase seen here. There was also a change in the origin of the published research. From 2014 to 2020, 2 out of 25 manuscripts had a primary author affiliation in China, compared with 292 out of 316 manuscripts between 2021 and 2024.

One hundred sixty-nine predictor variables were proposed as having statistically significant associations with the different health conditions reviewed in this meta-analysis. The three most commonly cited independent variables were the systematic immune-inflammation index, sleep health and serum vitamin D concentrations. The network analysis in Fig 4 illustrates the number of identified relationships and the complexity of the interactions, for the 16 health conditions in this systematic review with 4 or more associated predictors. The data underlying Fig 4 are aso shown in full in S1 Data Table B.

For the 138 health conditions studied in this meta-analysis, the number of studies proposing an association with an independent variable (biomarker or health/clinical indicator) ranged from 1 to 28; the mean average was 2 studies per condition, but the distribution was skewed, and most conditions were only associated to one independent variable. Some variables were included as both predictors and as outcomes, for example, C-reactive protein (CRP) was investigated as being associated with periodontitis (PMID 37481511), but in another study, a health behavior index was investigated as a predictor of elevated CRP levels (PMID 38628050). In the most extreme case, additional manuscripts could be generated by simply reversing dependent and independent variables in any statistical analysis, increasing the number of combinations of predictors and outcomes with no physiological justification or hypothesis. Depression was analyzed more frequently than any other condition, with 28 individual studies, all but 4 of which were published in 2023 or 2024. The associated independent variables are shown in Table 2 as a case study. The individual studies did not employ false discovery correction, but taken together, represent multiple hypotheses. To compensate for this, False Discovery Rate (FDR) correction using Benjamini-Yekutieli was then applied to these studies using a count of 28 potential relevant hypotheses. Of the 28 statistically significant associations, less than half (13) remained statistically significant after FDR correction.

Lines of inquiry this paper opens 13

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems perform peer review as effectively as humans? What gaps exist between benchmark performance and real deployment outcomes? What human oversight must AI research systems have? Do restrictions on reviewer LLM use actually shape peer review behavior? How do hallucinated citations emerge in AI scholarly output? What governance mechanisms can effectively constrain widely deployed AI systems?