AI Now Writes as Many Online Articles as Humans

Paper · Source
Expertise in the Age of AI Content

Source: Graphite · 2026-05-15

The number of articles published on the internet that are primarily AI-generated (50%) is equal to the number written by humans (50%).

ChatGPT launched in November 2022. Within the first 12 months, the percentage of primarily AI-generated articles jumped to 36%, and reached 48% by 24 months.

However, since Q1 2025 the percentage of primarily AI-generated articles has plateaued at roughly 50%. We previously published this finding with data up to May 2025, and new data confirms this trend.

Despite the prevalence of AI-generated articles on the web, we show in a separate study that these articles largely do not appear in Google and ChatGPT. We do not evaluate whether AI-generated articles get as much traffic as human-written articles, but we suspect that they do not.

Since ChatGPT launched in November 2022, many companies have explored publishing content generated by LLMs such as ChatGPT, Claude, and Gemini to grow their traffic across channels such as Google Search, social, and advertising. This is a cost-effective alternative to spending hundreds of dollars for humans to write content.

The quality of AI content is rapidly improving. In many cases, AI-generated content is as good or better than content written by humans (MIT Study). It is often hard for people to distinguish whether content is created by AI (Originality.ai Study).

We seek to evaluate the prevalence of AI-generated articles.

We observe significant growth in primarily AI-generated articles, coinciding with the launch of ChatGPT in November 2022. After only 12 months, primarily AI-generated articles accounted for 35.9% of articles published.

In Q1 2025, the quantity of primarily AI-generated articles being published on the web nearly equaled the quantity of human-written articles, 49.6% vs. 50.4%. In Q4 2025, primarily AI-generated articles surpassed human-written at 50.9%, before returning to 49.9% in Q1 2026.

While primarily AI-generated articles grew dramatically after ChatGPT launched, we do not see that trend continuing. Instead, the proportion of primarily AI-generated articles has remained relatively stable, near 50%, over the last five quarters. We hypothesize that this is because practitioners found that primarily AI-generated articles do not perform well in search, as shown in a separate study.

Accurate detection of AI-generated content is required to make claims about the prevalence of AI-generated articles on the web. There is considerable disagreement about the accuracy of AI detection algorithms, and many argue that detecting AI is impossible, or at best, highly inaccurate. Therefore, before classifying the articles in our data set, we evaluate the accuracy of the AI detectors.

All detectors have low false negative rates, especially for GPT-5, the most popular LLM as of May 2026.

Finally, we classify all 55.4k articles in our dataset using each detector to evaluate the percentage of articles that are primarily AI-generated. First, we compute the percentage of articles published in each quarter that are primarily AI-generated using each AI detector. Then, we simply take the average of those AI detector-level estimates.

We extended our Common Crawl sample to include articles published through March 2026.

We used three AI detectors instead of one, and averaged their detections. This method is preferable because we do not rely on the accuracy of a single detector. The methodology is similar in that for Pangram and Copyleaks, we consider an article primarily AI-generated when a majority of its content is detected as using AI. For GPTZero, we use its article-level predictions.

The overall story is the same: a steep rise in primarily AI-generated articles after ChatGPT's release and a plateau near 50% more recently. However, the percentage of primarily AI-generated articles we find by using multiple detectors is slightly lower than before (3.3 percentage points, on average), due to the more robust averaging method.

Many people incorporate AI into their content creation process. One strategy is to ask AI to create a first draft, then have a human in the loop to edit or rewrite it. We did not evaluate the accuracy of AI detectors using this strategy.

AI models continue to improve, and may become harder to detect. We only evaluate the false negative rate on articles generated by GPT-5, Gemini 3.1 Pro, and Claude Opus 4.6. The AI detection algorithm may have lower accuracy when applied to articles generated by other models.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can AI systems reliably guide voters without introducing political bias? Are AI-generated articles systematically disadvantaged in search ranking and user engagement? How reliably can humans and AI detectors identify machine-generated text? Why do confident AI outputs mislead human trust calibration? How does AI-generated content create social proof without authentic interaction? Can readers reliably distinguish AI-written text from human writing?