The AI Ghostwriter Effect: When Users Do Not Perceive Ownership of AI-Generated Text But Self-Declare as Authors

Paper · arXiv 2303.03283 · Published March 6, 2023
Expertise in the Age of AI Content

Human-AI interaction in text production increases complexity in authorship. In two empirical studies (n1 = 30 & n2 = 96), we investigate authorship and ownership in human-AI collaboration for personalized language generation. We show an AI Ghostwriter Effect: Users do not consider themselves the owners and authors of AI-generated text but refrain from publicly declaring AI authorship. Personalization of AI-generated texts did not impact the AI Ghostwriter Effect, and higher levels of participants’ influence on texts increased their sense of ownership. Participants were more likely to attribute ownership to supposedly human ghostwriters than AI ghostwriters, resulting in a higher ownership-authorship discrepancy for human ghostwriters. Rationalizations for authorship in AI ghostwriters and human ghostwriters were similar. We discuss how our findings relate to psychological ownership and human-AI interaction to lay the foundations for adapting authorship frameworks and user interfaces in AI in text-generation tasks.

Introduction. Imagine your short visit to New York is coming to an end, and rather than spending time writing a postcard yourself, you might ask GPT-3 to write a personalized postcard for you. Would you sign this postcard with your own name? The use of personalized artificial intelligence (AI) in human-AI collaboration has the potential to significantly impact the concept of authorship. One example of such collaboration can be seen in using language generation models, such as GPT-3, to assist with writing tasks. Previous HCI research has investigated user perspectives on collaborative creative writing [40, 71, 99], style transfer [88], text summarization [43], and perceived authorship with AI suggestions [72]. However, it has not yet investigated how people attribute and declare authorship for the generated text. Previous research in the social sciences has investigated the prevalence and rationalization of ghostwriting [18, 38, 89]. Ghostwriting, as a practice of using text produced by someone without crediting, is common in some academic fields, such as medicine [45, 108], but also in writing autobiographies or political speeches [14]. Reasons for the use of ghostwriters include Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s).

© 2023 Copyright held by the owner/author(s). Manuscript submitted to ACM arXiv:2303.03283v2 [cs.HC] 7 Nov 2023 Under Submission ’XX, 2023, Draxler et al. financial or political interests [45], a lack of time and writing expertise [14], and the pressure to obtain gratification, e.g., good grades [74]. Personalized Large Language Models (LLMs) are becoming accessible to the public and have the potential to become ubiquitous for everyday tasks that include text production. Therefore, it is necessary to understand how authorship is declared with personalized AI and to what extent AI is used similarly to human ghostwriters. This is particularly relevant when texts are personalized to users’ experiences and individual writing styles, and the lines between user and AI contributions start to blur.

In this paper, we approach human-AI interaction for personalized LLMs from an authorship perspective. Thus, we are interested in investigating the psychological processes involved when attributing ownership and authorship for text generation with personalized AI, i.e., with models fine-tuned to an individual user. Note that this is different from the focus in ongoing legal and ethical discussions of AI-generated material that will yield a prescriptive framework, i.e., how authorship should be declared, and that focuses much on non-personalized material in academic and professional contexts [91]. Consequently, we study authorship in a personal writing context, where declared authorship reflects individual decisions rather than explicit guidelines and where personalization to individual writing styles is an emerging practice.

To tackle the topic of ownership and authorship in personalized LLMs, we conducted two empirical studies that show the AI Ghostwriter Effect of personalized AI. We define this effect as the use of personalized generative AI without credit to the AI. Thus, the AI acts as a ghostwriter. Overall, we found that users do not judge themselves to be the authors of AI-generated text but often refrain from publicly declaring authorship of AI. In Study 1, we compared whether personalization and different interaction methods affect perceived ownership and declared authorship. In Study 2, we replicate the AI Ghostwriter Effect in a large sample and compare it to a human ghostwriter.

We show that personalization is not critical to the AI Ghostwriter Effect (Study 1). Subjective control over the interaction and the content increases the sense of ownership. Moreover, people use similar rationalizations for AI ghostwriters and human ghostwriters (Study 2, pre-registered at https://aspredicted.org/RKV_ZXX)1. From this, we motivate how to expand common authorship declaration frameworks (e.g., the CRediT taxonomy2, [1, 3, 13]) to account for the support of AI from a user-centered perspective. Understanding the interplay between control, ownership, and authorship also informs the interaction design of future AI-supported text-generation systems.

Related work. 2.1 Generative Models and Personalization In recent years, various generative models have emerged for tasks such as text-to-text [15] and text-to-image3 generation, music composition based on prompts4, and text simplification [111]. Other models generate images [90], poetry [102], rap lyrics [85], or music [57] in the style of previous works. Generative models often implement a Transformer architecture that includes an encoder and a decoder, utilizing an attention mechanism that allows back-references to The AI Ghostwriter Effect Under Submission ’XX, 2023, prior content [105]. These models are typically trained to predict probable next tokens from the previous text. Thus, they learn connections between subsequent words and are well-suited for text generation.

In this work, we focus on text generation with GPT-3. Released in 2020, GPT-3 currently is a particularly large and powerful language model [25, 27]. It uses deep learning methods such as attention mechanisms and the Transformer architecture with autoregressive pretraining [15]. This architecture has improved the generation of long coherent text [103] because of the model’s ability to give special attention to key textual features and to connect interrelated words over longer passages [32]. GPT-3 produces text5 that statistically fits well with a given natural language prompt, using its syntactic ability to associate words without understanding the semantics and context of the query [34]. The text quality is often comparable to human-written texts in terms of grammar, coherency, and fluency, but it can also produce nonsensical or incorrect content [32, 103]. Like many other machine learning systems, GPT-3 may also reproduce biases found in its training data, such as racial or gender biases [15, 76]. One use case in which GPT-3 performs well is “few-shot learning,” where demonstrations of input-output pairs for the task are included with the prompt [15]. The choice of these in-context examples has a significant influence on the content and the quality of the generated text [75].

The same idea can also be used for personalization: the output of the system is conditioned by the information available about the user [111] and their context [29] to better match generated texts to a writer’s expectations and needs. According to Wang et al. [106], personalization can happen at two levels: it can include factual knowledge such as personal data, user attributes, and user preferences. The second category is defined as stylistic modifiers, which can either be situational or personal. For example, Gmail Smart Compose interpolates between a global and a personal model trained on individual sent messages [19].

Method. 3 RESEARCH QUESTIONS AND HYPOTHESES We pose the following research questions:

RQ1: Does the sense of ownership match the declaration of authorship for personalized AI-generated texts?

RQ2: Does the level of influence in human-AI interaction affect the sense of ownership?

RQ3: Does the sense of ownership depend on the quality of personalization?

RQ4: In what ways does the AI Ghostwriter Effect differ from human ghostwriters?

Our paper aims to investigate what affects authorship attribution for texts generated with personalized LLMs from an HCI perspective. First, we identify the psychological processes involved when attributing authorship in human-AI interaction. Note that we consider the process of attributing authorship (from the sense of ownership to the declaration of authorship), the perceived degree of control in human-AI interaction, the subjective quality of personalization, and the feeling of leading the human-AI interaction. In other words, the sense of ownership addresses the subjective side: does a person feel that they are the author of a text and that they own it? This also encompasses a sense of control over the product, as the two concepts are tightly linked (cf. Section 2.4). The declaration of authorship, on the other hand, refers to the entity a text is attributed to, e.g., in the header or byline.

If the AI system, here GPT-3, is used much like a human ghostwriter, then we hypothesize that there is an AI Ghostwriter Effect:

H1.1: participants will consider the AI (and not themselves) to be the owner of personalized AI-generated texts but H1.2: they will not declare AI authorship or AI support when publishing personalized AI-generated texts.

Second, we want to investigate how common interaction methods affect the process of attributing authorship. This is important as designers of interfaces are faced with a large body of normative statements on authorship with little knowledge of how user interfaces affect the declaration of authorship and what user experience they create during human-AI interaction.

To effectively design interactive LLM-based applications with authorship in mind, we have to know how different interaction methods affect authorship for personalized AI-generated texts. Lehmann et al. [72] found that the sense of authorship positively correlates with the degree of influence over the AI contribution. Here, objective control refers to the degree of influence that users have over the AI-generated text, e.g., by employing interaction methods categorized as Writing, Editing, Choosing, and Getting, while perceived control refers to the subjective level of being able to The AI Ghostwriter Effect Under Submission ’XX, 2023, affect the generated text a priori. Leadership, in turn, refers to the user’s perceived initiative to shape the outcome, i.e., agency during the interaction.

Thus, we also assess perceived control and leadership, and we hypothesize that H2.1: the degree of influence over the generated text, as realized by different interaction methods, affects the sense of control, and H2.2: the degree of influence over the generated text affects the sense of leadership.

Third, we investigate how AI personalization affects this process. Here, personalization refers to the customization of the AI model using fine-tuning to suit an individual user’s specific preferences and writing style. We explicitly investigate the case of personalized text generation because, in this case, the lines between the AI model and the user blur. However, recent studies have found that usability and user experience do not depend on the quality of personalization and adaptation for fine-tuned AI systems [104]. Even non-adaptive systems are deemed adaptive and personalized if introduced as such [68]. If this holds for personalized LLMs, merely labeling the AI as personalized (i.e., placebo-personalization) should yield comparable effects on the sense of ownership.

We consider two levels of personalization quality – personalization by fine-tuning and placebo-personalization – and we hypothesize that H3: the sense of ownership is independent of the quality of personalization for the AI model.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How should human-AI contributions be measured, disclosed, and verified? How do writers navigate authorship and delegation with AI? How reliably can humans and AI detectors identify machine-generated text? Can readers reliably distinguish AI-written text from human writing? Why do confident AI outputs mislead human trust calibration? How does AI-generated content create social proof without authentic interaction? Does AI assistance erode cognitive skills while inflating perceived competence? Are AI-generated articles systematically disadvantaged in search ranking and user engagement? Does disclosing AI authorship change how audiences evaluate the writing?