AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances

Paper · arXiv 2409.11360 · Published September 17, 2024
Expertise in the Age of AI Content

ABSTRACT Large language models (LLMs) are being increasingly integrated into everyday products and services, such as coding tools and writing assistants. As these embedded AI applications are deployed globally, there is a growing concern that the AI models underlying these applications prioritize Western values. This paper investigates what happens when a Western-centric AI model provides writing suggestions to users from a different cultural background. We conducted a cross-cultural controlled experiment with 118 participants from India and the United States who completed culturally grounded writing tasks with and without AI suggestions. Our analysis reveals that AI provided greater efficiency gains for Americans compared to Indians. Moreover, AI suggestions led Indian participants to adopt Western writing styles, altering not just what is written but also how it is written. These findings show that Western-centric AI models homogenize writing toward Western norms, diminishing nuances that differentiate cultural expression.

Introduction. Large language models (LLMs) are increasingly shaping online discourse and are being integrated into everyday applications to improve user productivity and experience. These include chatbots that aid in content comprehension1, code completion systems that streamline programming2, and autocomplete tools that facilitate faster writing [16]. Writing, in particular, has become a prominent area where LLMs offer substantial assistance, supporting tasks such as scientific writing [30], storytelling [79], and journalism [65]. These applications provide inline writing suggestions, which users can accept or reject in real-time. Such AI writing suggestions have become integral to email clients (e.g., Gmail Smart Compose [16]), note-taking applications (e.g., Notion), and word processors (e.g., Google Docs). Their growing utility is evident from their integration into popular web browsers like Google Chrome and Microsoft Edge, which now offer autocomplete as a native feature, enabled by default3. While LLM-embedded applications like text autocomplete have gained widespread global adoption, these technologies also carry significant risks, such as perpetuating harmful stereotypes about marginalized groups and underserved communities. For instance, LLMs have been shown to depict Muslims as terrorists [1] and reinforce negative stereotypes about disabled individuals [27, 85]. Additionally, research has revealed that LLMs prioritize Western norms and values in their interactions [14, 44], resulting in representational harms for diverse non-Western cultures [9, 68, 77]. While explicit cultural stereotyping shown by prior work is deeply problematic, an even more insidious issue lies in the subtle biases LLMs can introduce through their suggestions. For example, recent work shows that autocomplete suggestions can change users’ language [40], writing [66], and even attitudes about social topics [43, 86]. Furthermore, unlike chat-based applications such as ChatGPT, LLM-embedded applications do not allow users to prompt and fine-tune the model to suit their cultural preferences. This creates situations where users and AI may have different cultural values, potentially leading to conflicts over whose norms should be expressed. While prior studies have explored cultural harms in LLMs using open-ended prompts, no research has yet examined such cultural clashes in embedded AI applications. In this paper, we examine how AI suggestions influence user-generated content when the AI and users share the same cultural identity versus when they do not. In particular, we ask:

RQ1: Does writing with a Western-centric AI provide greater benefits to users from Western cultures, compared to those from non-Western cultures? RQ2: Does writing with a Western-centric AI homogenize the writing styles of non-Western users toward Western styles?

To answer these questions, we conducted a cross-cultural experiment with 118 users from India and the US recruited through Prolific, an online crowdsourcing platform. Participants from both cultures were asked to complete writing tasks in English. We designed the task prompts using Hofstede’s Cultural Onion framework [39], which allowed us to elicit cultural practices ranging from explicit (e.g., food) to implicit (e.g., rituals). Each participant was randomly assigned either to the AI condition receiving autocomplete suggestions from GPT-4o while writing or the No AI condition writing organically without AI assistance. Subsequently, we compared the essays written by participants in the four experimental groups (Indian and American users writing with and without AI suggestions). Our analysis yielded two key results. In response to RQ1, we found that while AI boosts productivity for both Indian and American participants, the gains are higher for American participants. This discrepancy raises concerns about quality-of-service harms [77] for non-Western users, wherein they need to put in more effort to achieve similar benefits. In response to RQ2, we found that AI caused Indian participants to write more like Americans, thereby homogenizing writing toward Western styles and diminishing nuances that differentiate cultural expression. Worryingly, AI influences not just what is written (e.g., shifting preferences toward Western cultural artifacts such as food items), but also more ingrained elements of how it’s written, silently erasing non-Western styles of cultural expression. For instance, with AI suggestions, Indians lose cultural nuance and describe their own food and festivals from a Western gaze. We discuss the implications of these cultural harms from a colonial lens and propose strategies to mitigate them. Overall, we make the following contributions to the HCI and AI communities:

Related work. We first review scholarly work exploring cultural bias in AI models, emphasizing how LLMs prioritize Western norms and values. Given our focus on AI-based writing suggestions, we then situate our work within the emerging HCI scholarship on designing, developing, and evaluating AI technologies for writing support.

An emerging body of HCI and AI scholarship has examined the cultural composition of generative AI technologies [2], and found that these models center Western norms and values [14, 83]. For example, Qadri et al. [68] showed that text-to-image models fail to generate cultural artifacts, amplify hegemonic cultural defaults, and perpetuate cultural tropes when depicting non-Western cultures. Even within the West, these AI models tend to prioritize US-centric values and artifacts [9, 44, 74, 85]. Furthermore, LLMs trained and probed in non-Western languages (e.g., Arabic) continue to show Western bias [6, 61]. Recognizing these biases, recent work has proposed benchmark datasets to evaluate the multicultural knowledge of LLMs [17, 70]. Together, these studies reveal the harms of culturally incongruent AI models [67]. They underscore how data sourced from low-paid workers in non-Western regions is commodified to create models that inadequately serve these regions, a phenomenon known as data colonization [20]. Moreover, these systems reinforce Western cultural hegemony, perpetuate cultural erasure, and amplify stereotypes [68], echoing the dynamics of colonial times and giving rise to AI colonialism. [34, 81]. Addressing these cultural biases is crucial not only to prevent representational, allocative, and quality-of-service harms [4, 8, 77] but also to avoid culture clashes that arise when the values embedded in AI models diverge from those of their users [36, 67]. Early work has acknowledged these challenges of value plurality, questioning what it means to imbue AI with human values when those values vary across cultures [82], and which values a model should prioritize when producing a singular output [44]. As generative AI technologies like LLMs are increasingly integrated into global products and services, these questions have become critically important. The commercial growth has also driven a shift toward embedded AI applications, where the underlying models are obscured from users. This opacity makes it challenging to mitigate cultural biases through open-ended prompting or fine-tuning, thereby surfacing cultural clashes. We examine this culture clash that arises when users interact with embedded AI models that are culturally distant from them. To our knowledge, this is the first to reveal cultural homogenization effects when embedded AI features continue to be powered by culturally biased models.

Method. We conducted a controlled experiment with 118 participants, comprising 60 Indian and 58 American users, recruited through the online crowdsourcing platform Prolific. Participants completed short writing tasks designed to elicit cultural values and artifacts. We had a standard 2 × 2 study design in which participants from India and the US completed the tasks with or without AI suggestions. Subsequently, we compared the essays written by participants from the four groups. Below, we describe our methodology in detail.

To examine the impact of writing with a culturally incongruent model, we simulated a “cultural distance” between the users and the model. Since AI models are often aligned with Western cultures and values [9, 44, 59, 68], we conducted the study with participants from the US (representing a smaller cultural distance) and India (representing a larger cultural distance). Participants from both cultures were randomly assigned to complete writing tasks with or without AI suggestions, resulting in a standard 2 × 2 betweensubjects study design. Hence, each participant was assigned to one of four experimental groups: Indians writing with or without AI, and Americans writing with or without AI4. These groups are summarized in Table 1. The two control groups (without AI) helped us capture the natural writing styles of participants from both cultures, forming a baseline for comparison with the treatment groups. The treatment groups (with AI) were designed to investigate whether writing styles changed when AI suggestions were introduced.

Each participant was required to complete four writing tasks in English. Although English is not a native language for most Indians, we expected our participants to be fluent due to being on Prolific, a platform run entirely in English. Indeed, 59 of our 60 Indian participants reported speaking English. The task topics were designed to elicit various aspects of culture defined by Hofstede in his “Cultural Onion” [39] (Figure 2). Hofstede’s conceptualization of culture, such as the six cultural dimensions [37], and the cultural onion [39], has been widely used in prior work to operationalize culture [6, 32, 54]. In particular, the cultural onion framework uses the metaphor of a layered onion to model culture, envisioning an outsider progressing through layers of explicit cultural practices (symbols, heroes, rituals) to ultimately grasp the core of the culture, its implicit values. This layered approach allows us to elicit both explicit cultural practices and implicit values. Specifically, symbols are words and objects that hold a certain meaning in a culture, including language, art, clothes, food, etc. Symbols are not fixed and evolve but are observable by outsiders to the culture. Heroes are people idolized by a culture, often because they possess characteristics that are valued within the culture. They may be real (e.g., APJ Abdul Kalam in India) or fictional (e.g., Rocky Balboa in the US). Rituals are collective activities deemed important by the people of the culture. This may include holidays and festivals, ways of greeting, or other social customs (e.g., sauna). These three layers are collectively referred to as practices and can be observed and practiced by outsiders even without necessarily understanding their underlying cultural meanings. Finally, at the core of the onion are the values, which are preferences for certain states of affairs (e.g., individualism or collectivism). They are static and so ingrained that even people within the culture may be unaware of them. To holistically evoke these cultural nuances, we designed four writing tasks for our study, one for each layer of the cultural onion. The task prompts are shown in Table 2. To select these prompts, we chose representative artifacts from the outer layers of the onion and designed a relatable digital interaction (writing an email to a superior) to evoke underlying cultural values. There was an additional task in the study to function as an attention check. Attention checks are commonly used tools to improve the quality of data obtained from crowdsourcing platforms [35, 47].

Discussion. Our results show that AI-powered writing suggestions yield greater benefits for American users compared to Indian users. Moreover, AI suggestions homogenize writing styles between Indians and Americans, pushing Indian users to adopt American writing styles and diminishing cultural nuances in writing. Worryingly, the changes induced by AI suggestions are both explicit (e.g., changing cultural preferences), as well as subtle and deeply ingrained (e.g., lexical diversity, exoticization, Westernization), making it harder to identify the cultural erasure they might be causing. In the following sections, we discuss cultural differences in AI attitudes, examine the implications of the observed homogenization of writing styles, and propose strategies to mitigate the harms of cultural homogenization and AI-driven neocolonialism.

As we showed in Section 4.1, Indian participants exhibited higher engagement with AI suggestions: a larger proportion of their essays were composed of AI-generated text, and they accepted more AI suggestions compared to Americans. From a purely causal lens, this difference in AI engagement could be seen as a confounding factor, potentially undermining our findings. For example, it could suggest that the homogenization effects are simply a result of Indians’ higher AI usage. However, we argue that this is not a flaw, but an important finding of our study. To fully understand the differential AI engagement seen in our study, we turn to HCI scholarship for a socio-technical view of human-AI interaction. The way people perceive the benefits and harms of emerging technologies varies across cultures and geographic regions [53]. For instance, people in individualistic cultures (e.g., the US) tend to evaluate new technologies through direct and formal sources, while those in collectivistic cultures (e.g., India) rely more on peer feedback [51]. AI technologies are no different—they are cultural artifacts [22] and perceptions about them vary across cultures, both at the national [23] and user levels [53]. Research shows that users in non-Western settings express more positive emotions toward conversational agents, whereas users in the US tend to be more ambivalent [53]. Non-Western users also show higher levels of trust in AI technologies than their American counterparts [78]. In writing contexts, particularly when composing in English, AI-based writing suggestions provide more utility to non-native speakers than to native speakers, leading them to engage more with the suggestions [13]. This positive attitude often manifests as overreliance on AI [12], especially among novice users in non-Western settings [62]. These cultural differences in AI overreliance may be exacerbated as embedded AI applications reach more diverse users. While past research has shown that overreliance can be mitigated with AI explanations [80, 84], this approach becomes less viable in new-age interactive and embedded AI tools, where providing explanations may not be feasible. For instance, in writing applications, presenting explanations for every suggestion may disrupt the user’s flow, hindering productivity. Furthermore, our study shows that embedded AI applications promising productivity gains may inadvertently push non-Western users toward overreliance by delivering diminished benefits to them. We argue that higher trust and AI reliance are cultural characteristics that significantly influence human-AI interaction. It is important to acknowledge these cross-cultural differences in trust, adoption, and degree of AI reliance. By trying to eliminate them in our analyses as mere confounders, we may fail to notice the cultural harms they are causing, risking a future where cultural nuances are erased.

Conclusion. We conducted a cross-cultural controlled experiment with 118 Indian and American participants to examine the consequences of interacting with a culturally different model. Participants completed writing tasks designed to elicit cultural practices and values. They were randomly assigned to one of two conditions: the AI condition, where they received inline AI suggestions to assist with their writing, or the No AI condition, in which they wrote without any AI support. Our analysis of writing logs and essays revealed two key findings. First, AI suggestions offered greater productivity gains to American users, leaving non-Western users to put in more effort to achieve similar benefits. Second, AI homogenized writing styles towards Western norms, subtly influencing Indians to adopt Western writing patterns. This homogenization provides concrete evidence of AI colonialism, where these models reinforce Western cultural hegemony. It is crucial to address these biases before culturally biased content proliferates and becomes part of the training data for future AI models.

Limitations. 5.4 Limitations and Future Work The terms “Western” and “non-Western” are broad cultural labels, as US culture does not fully represent all Western cultures, nor does Indian culture encompass all non-Western cultures. Even within “India” and “the US,” there exist rich and diverse subcultures. In using these categorizations, we draw upon prior work [9, 68] that has used geographical boundaries as a proxy for cultural differences. However, future work is required to determine if our results generalize to other countries and sub-cultures. In the wake of our findings, this is especially important to understand the scale at which cultural erasure might be happening, especially for cultures that are already marginalized. Ironically, the community’s focus on broad cultural categorization (e.g., treating “India” or “Africa” as monolithic cultures) may itself contribute to cultural erasure.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do writers navigate authorship and delegation with AI? Can readers reliably distinguish AI-written text from human writing? How reliably can humans and AI detectors identify machine-generated text? How does AI-generated content create social proof without authentic interaction? How do educators verify student capability when AI can produce indistinguishable work?