Are Large Language Models a Threat to Digital Public Goods? Evidence from Activity on Stack Overflow

Paper · arXiv 2307.07367 · Published July 14, 2023
Expertise in the Age of AI Content

Abstract Large language models like ChatGPT efficiently provide users with information about various topics, presenting a potential substitute for searching the web and asking people for help online. But since users interact privately with the model, these models may drastically reduce the amount of publicly available human-generated data and knowledge resources. This substitution can present a significant problem in securing training data for future models. In this work, we investigate how the release of ChatGPT changed human-generated open data on the web by analyzing the activity on Stack Overflow, the leading online Q&A platform for computer programming. We find that relative to its Russian and Chinese counterparts, where access to ChatGPT is limited, and to similar forums for mathematics, where ChatGPT is less capable, activity on Stack Overflow significantly decreased. A difference-in-differences model estimates a 16% decrease in weekly posts on Stack Overflow. This effect increases in magnitude over time, and is larger for posts related to the most widely used programming languages. Posts made after ChatGPT get similar voting scores than before, suggesting that ChatGPT is not merely displacing duplicate or low-quality content. These results suggest that more users are adopting large language models to answer questions and they are better substitutes for Stack Overflow for languages for which they have more training data. Using models like ChatGPT may be more efficient for solving certain programming problems, but its widespread adoption and the resulting shift away from public exchange on the web will limit the open data people and models can learn from in the future.

Introduction. Over the last thirty years, humans have constructed a vast library of information on the web. Using powerful search engines anyone with an internet connection can access valuable information from online knowledge repositories like Wikipedia, Stack Overflow, and Reddit. New content and discussions posted online are quickly integrated into this ever-growing ecosystem, becoming digital public goods used by people all around the world to learn new technologies and solve their problems (Hess and Ostrom, 2003; Henzinger and Lawrence, 2004; Lemmerich et al., 2019; Piccardi et al., 2021).

More recently, these public goods have been used to train artificial intelligence (AI) systems, in particular, large language models (LLMs) (Vaswani et al., 2017). For example, the LLM ChatGPT (OpenAI, 2023) answers user questions by summarizing the information contained in these repositories.

The remarkable effectiveness of ChatGPT is reflected in its quick adoption (Teubner et al., 2023) and application across diverse fields including auditing (Gu et al., 2023), astronomy (Smith and Geach, 2023), medicine (Kanjee et al., 2023), and chemistry (Guo et al., 2023). Randomized control trials show that using LLMs significantly boosts productivity in computer programming, professional writing, and customer support tasks (Peng et al., 2023; Noy and Zhang, 2023; Brynjolfsson et al., 2023). Indeed, the widely reported successes of LLMs like ChatGPT suggest that we will observe a significant change in how people search for, create and share information online.

Ironically, if LLMs like ChatGPT present substitute traditional ways of searching and interrogating the web, then they will displace the very human behavior that generated their original training data.

User interactions with ChatGPT are the exclusive property of OpenAI, its creator. Only OpenAI will be able to learn from the information contained in these interactions. As people begin to use LLMs instead of online knowledge repositories to find information, contributions to these repositories will likely decrease, diminishing the quantity and quality of these digital public goods. While such a shift would have significant social and economic implications, we have little evidence on whether people are actually substituting their consumption and creation of digital public goods with ChatGPT.

The aim of this paper is to evaluate the impact of LLMs on the generation of open data on question- and-answer (Q&A) platforms. Since LLMs perform relatively well on software programming tasks (Peng et al., 2023), we study Stack Overflow, the largest online Q&A platform for software development and programming. We present three results. First, we examine whether the release of ChatGPT has decreased the volume of posts, i.e. questions and answers, posted on the platform. We measure the overall effect of ChatGPT’s release on Stack Overflow activity using a difference-in-differences model. We compare the weekly posting activity on Stack Overflow against that of four comparable Q&A platforms. These counterfactual platforms are less likely to be affected by ChatGPT either because their users are less able to access ChatGPT or because ChatGPT performs poorly in questions discussed on those platforms. We find that posting activity on Stack Overflow decreased by about 16% following the release of ChatGPT, increasing over time to around 25% within six months.

Second, we investigate whether ChatGPT is simply displacing simpler or lower quality posts on Stack Overflow. To do so, we use data on up- and downvotes, simple forms of social feedback provided by other users to rate posts. We observe no change in the votes posts receive on Stack Overflow since the release of ChatGPT. This finding suggests that ChatGPT is displacing a wide variety of Stack Overflow posts, including high-quality content.

Third, we study the heterogeneity of the impact of ChatGPT across different programming languages discussed on Stack Overflow. We test for these heterogeneities using an event study design. We observe that posting activity in some languages like Python and Javascript has decreased significantly more than the global site average. Using data on programming language popularity on GitHub, we find that the most widely used languages tend to have larger relative declines in posting activity.

Our analysis points to several significant implications for the sustainability of the current AI ecosystem. The first is that the decreased production of open data will limit the training of future models (Villalobos et al., 2022). LLM-generated content itself is an ineffective substitute for training data generated by humans for the purpose of training new models (Gudibande et al., 2023; Shumailov et al., 2023; Alemohammad et al., 2023). One analogy is that training an LLM on LLM-generated content is like making a photocopy of a photocopy, providing successively less satisfying results (Chiang, 2023).

Method. 2.1 Stack Exchange and Segmentfault data To understand the effect ChatGPT can have on digital public goods, we compare the change in Stack Overflow’s activity with the activity on a set of similar platforms. These platforms are similar to Stack Overflow in that they are technical Q&A platforms, but are less prone to substitution by ChatGPT given their focus or target group. Specifically, we focus on the Stack Exchange platforms Mathematics and Math Overflow and on the Russian-language version of Stack Overflow. We also examine a Chinese-language Q&A platform on computer programming called Segmentfault.

Mathematics and Math Overflow focus on university- and research-level mathematics questions respectively. We consider these sites to be less susceptible to replacement by ChatGPT given that, during our study’s period of observation, the free-tier version of ChatGPT performed poorly (0-20th percentile) on advanced high-school mathematics exams (OpenAI, 2023), and was therefore unlikely to serve as a suitable alternative to these platforms.

The Russian Stack Overflow and the Chinese Segmentfault have the same scope as Stack Overflow, but target users located in Russia and China, respectively. We consider these platforms to be less affected by ChatGPT given that ChatGPT is officially unavailable in the Russian Federation, Belarus, Russianoccupied Ukrainian territory, and the People’s Republic of China. Although people in these places can and do access ChatGPT via VPNs (Kreitmeir and Raschky, 2023), such barriers still represent a hurdle to widespread fast adoption.

We extract all posts (questions or answers) on Stack Overflow, Mathematics, Math Overflow, and Russian Stack Overflow from their launch to early June 2023 using https://archive.org/details/ stackexchange. We scraped the data from Segmentfault directly. Our dataset comprises 58 million posts on Stack Overflow, over 900 thousand posts for the Russian-language version of Stack Overflow, 3.5 million posts on Mathematics Stack Exchange, 300 thousand posts for Math Overflow, and about 300 thousand for Segmentfault. We focus our analysis on data from January 2019 to June 2023, noting that our findings are robust to alternative time windows.

For each post, our dataset includes the number of votes (up – positive feedback, or down – negative feedback) the post received, the author (user), and whether the post is a question or an answer. Furthermore, each post can have up to 5 tags – predefined labels that summarize the content of the post, for instance, an associated programming language. For more details on the data used, we refer the reader to section 5. From this point forward, we will refer to Mathematics, Math Overflow, Russian Stack Overflow, and Segmentfault, along with their corresponding posts, as the counterfactual platforms and posts.

2.2 Models Difference-in-differences We estimate the effect of ChatGPT for posting activity on Stack Overflow using a difference-in-differences method with four counterfactual platforms. We aggregate posting data at platform- and week-level and fit a regression model using ordinary least squares (OLS):

IHS(Postsp,t) = αp + λt + β × Treatedp,t + X p∈P θpt + εp,t (1) where Postsp,t is the number of posts on platform p in a week t, which we transform using the inversehyperbolic sine function (IHS) (Burbidge et al., 1988).1 αp are platform fixed effects, λt are time (week) fixed effects, θp are platform-specific linear time trends, and εp,t is the error term.

The coefficient of interest is β, which captures the estimated effect of ChatGPT on posting activity on Stack Overflow relative to the less affected platforms: Treated equals one for weeks after the release of ChatGPT (starting November 27, 2022) when the platform p is Stack Overflow and zero otherwise. We report robust standard errors clustered at the monthly level.

Discussion. The rate at which people have adopted ChatGPT is one of the fastest in the history of technology (Teubner et al., 2023). It is essential that we better understand what activities this new technology displaces and what second-order effects this substitution may have (Schumpeter, 1942; Aghion and Howitt, 1992). This paper shows that after the introduction of ChatGPT there was a sharp decrease in human content creation on Stack Overflow. We compare the decrease in activity on Stack Overflow with other Stack Exchange platforms where current LLMs are less likely to be used. Using a difference-in-differences model, we find about 16% relative decrease in posting activity on Stack Overflow, with a larger effect in later months.

We observed no large change in social feedback on posts, measured using votes, following ChatGPT’s release, suggesting that average post quality has not changed. Posting activity related to more popular programming languages decreased more on average than that for more niche languages. These results suggest that users partially substituted Stack Overflow with ChatGPT. Consequently, the wide adoption of LLMs can decrease the provision of digital public goods, in particular, the open data previously generated by interactions on the web.

Despite these shortcomings, our results have important implications for the future of digital public goods. Before the introduction of ChatGPT, more human-generated content was posted to Stack Overflow, forming a collective digital public good due to their non-rivalrous and non-exclusionary nature – anyone with internet access can view, absorb, and extend this information, without diminishing the value of the knowledge. Now, this information is rather fed into privately owned LLMs like ChatGPT. This represents a significant and trending shift of knowledge from the public domain to the private ones.

This observed substitution effect poses several issues for the future of artificial intelligence in general.

The first is that if language models crowd out open data creation, they will be limiting their own future training data and effectiveness. The second is that owners of the current leading models have exclusive access to user inputs and feedback, which, with a relatively smaller pool of open data, gives them a significant advantage against new competitors in training future models. Third, the decline of public resources on the web would reverse progress made by the web toward democratizing access to knowledge and information. Finally, the consolidation of humans searching for information around one or a few language models could narrow our explorations and focus our attention on mainstream topics. We briefly elaborate on these points, then conclude with a wider appeal for more research on the political economy of open data and AI, and how we can incentivize continued contributions to digital public goods.

Training future models Our findings suggest that the widespread adoption of ChatGPT may make it difficult to train few iterations (Taleb, 2012). Though researchers have already expressed concerns about running out of data for training AI models (Villalobos et al., 2022), our results show that the use of LLMs can slow down the creation of new data. Given the growing evidence that data generated by LLMs cannot effectively train new LLMs (Gudibande et al., 2023; Shumailov et al., 2023; Alemohammad et al., 2023), modelers face the real problem of running out of useful data. If ChatGPT truly is a “blurry JPEG” of the web (Chiang, 2023), then, in the long run, it cannot effectively replace its most important input: data derived from human activity.

Limitations. Our results and data have some shortcomings that point to open questions about the use and impact of LLMs. First, while we can present strong evidence that ChatGPT decreased the posting activity in Stack Overflow, we can only partially assess quality of posting activity using data on upvotes and downvotes. Users may be posting more challenging questions, ones that LLMs cannot (yet) address, to Stack Overflow. Future work should examine whether continued activity on Stack Overflow is more complex or sophisticated on average than posts from prior to ChatGPT release. Similarly, ChatGPT may have reduced the volume of duplicate questions about simple topics, though this is unlikely to impact our main results as duplicates are estimated to account for only 3% of posts (Correa and Sureka, 2013), and we do not observe significant changes in voting outcomes.

A second limitation of our work is that we cannot observe the extent to which Russian- and Chineselanguage users of the corresponding Q&A platforms are actually hindered from accessing ChatGPT; indeed recent work has shown a spike in VPN and Tor activity following the blocking of ChatGPT in Italy (Kreitmeir and Raschky, 2023).

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can AI systems reliably guide voters without introducing political bias? Are AI-generated articles systematically disadvantaged in search ranking and user engagement? How should retrieval strategies adapt to multi-step reasoning demands? Can AI chatbots provide mental health support without reinforcing harmful beliefs? How does AI adoption reshape collaboration patterns in knowledge work? How do network effects and self-selection distort aggregated rating accuracy? How do AI hiring systems affect authenticity, fairness, and candidate preferences?