Undermining Mental Proof: How AI Can Make Cooperation Harder by Making Thinking Easier

Paper · arXiv 2407.14452 · Published July 19, 2024
Expertise in the Age of AI Content

Large language models and other highly capable AI systems ease the burdens of deciding what to say or do, but this very ease can undermine the effectiveness of our actions in social contexts. We explain this apparent tension by introducing the integrative theoretical concept of “mental proof,” which occurs when observable actions are used to certify unobservable mental facts. From hiring to dating, mental proofs enable people to credibly communicate values, intentions, states of knowledge, and other private features of their minds to one another in low-trust environments where honesty cannot be easily enforced. Drawing on results from economics, theoretical biology, and computer science, we describe the core theoretical mechanisms that enable people to effect mental proofs. An analysis of these mechanisms clarifies when and how artificial intelligence can make low-trust cooperation harder despite making thinking easier.

Introduction. The widespread availability of generative artificial intelligence means that anyone can now cheaply and convincingly simulate the output of human mental effort across an unprecedented variety of tasks. This promises numerous benefits across nearly every aspect of society, but it also has begun to disrupt an equally broad set of social practices, such as sincere apologies (Glikson and Asscher, 2023), college assessment (Fitria, 2023; Cardon et al., 2023), online dating (Wu and Kelly, 2020), and wedding vows (LaGorce, 2023). In light of these developments, many have come to see the technology as a doubleedged sword: artificial intelligence cuts the cost of thinking, but it also—and, as we will argue, for that very reason—threatens vital elements of the existing social fabric, such as trust (Glikson and Woolley, 2020), privacy (Jain et al., 2023), public safety (Leslie, 2019), and even democracy itself (Allen and Weyl, 2024; Jungherr, 2023). Indeed, a majority of survey respondents now feel “more concerned than excited” about the increased use of artificial intelligence in their daily lives (Tyson, 2023). Despite the urgency of these concerns, the scientific community has struggled to articulate general principles that explain why lowering mental costs through artificial intelligence undermines such a wide variety of seeming unrelated social practices. We may “know it when we see it” in any particular case, but the lack of integrative scientific frameworks has made it difficult to formulate general solutions to existing problems or foresee future harms. In this paper, we highlight the key role that “mental proof” plays in facilitating cooperation in low-trust environments. Mental proofs are observable actions taken to certify unobservable facts about the minds who perform them. As we describe fully below, people use two distinct mechanisms to substantiate mental proofs: signaling theory (an idea primarily studied in economics and biology) and proof of knowledge protocols (studied in computer science). To function, both mechanisms rely on implicit assumptions about the cost structure of organic mental activity—assumptions that are rapidly being disrupted by the proliferation of artificial intelligence in daily life. An appreciation for the structure of mental proofs therefore helps elucidate the underlying logic of artificial intelligence’s various social consequences. In highlighting the importance of mental proof and its relationship to thinking machines, our paper contributes to a wider, cross-disciplinary attempt to proactively understand and address the technology’s various social consequences (e.g., Solaiman et al., 2023; Weidinger et al., 2021; Mirsky and Lee, 2021) We illustrate the practical importance of mental proof in two “worked examples” drawn from everyday social domains: sincere apology and subculture formation. In both cases, we discuss how low-cost simulations of intelligent behavior undermine the efficacy of mental proofs and the vital social benefits they provide. As our discussion makes clear, mental proof is most valuable in situations where honesty cannot easily be enforced. This implies that the category of harms we discuss will disproportionately impact those who are not already embedded in high-trust networks and formal institutions, thereby reinforcing existing structural barriers to social mobility and economic development. Our analysis does, however, suggest effective strategies for mitigating these deleterious consequences, and we conclude with a sketch of the framework’s implications for policy, technology, and everyday life.

Method. Communication can greatly enhance coordination. Our species’ remarkable capacity for coordination relies, at its foundation, upon an ability to reliably externalize the nuances of our internal mental states—not just our beliefs, but also our intentions, values, preferences, skills, understandings, commitments, abilities, etc.—in ways that others will not only understand, but trust. As has been pointed out by both behavioral scientists (e.g., Tomasello et al., 2005) and philosophers of mind (e.g, Gilbert, 1990; Bratman, 1992), people’s ability to share and understand intentions, in particular, is essential to the formation of collaborative acts that range from going on a walk together to drafting a new constitution. Sometimes this is easy: in many contexts, we can simply speak our minds and others will have good reason to believe us. Despite its many conveniences, however, “cheap talk” breaks down when people have incentives to lie (Farrell, 1987). In such contexts, mere assertions lose their credibility, and even those who try to tell the truth will be dismissed. A variety of social institutions help enforce honesty and thereby preserve the benefits of its coordinating function, most notably reputation (Fehr, 2009), norms (Elster, 1989), and formal punishment (North, 1991). Unfortunately, however, these institutions are not always available—e.g., before reputation is established, when claims are difficult to verify, or in places where cultural and legal institutions are weak. In these “low-trust” contexts, interacting parties must not only state, but endeavor to prove, claims about their minds. One way people furnish such proof is by taking observable actions which (given certain assumptions about the capabilities and structure of human brains) provide strong grounds for a relevant claim about the mind. We refer to such behaviors as constituting mental proof. The validity of mental proofs are primarily underwritten by two separate mechanisms: signaling theory and proof of knowledge protocols. We review each, in turn.

Signaling theory was introduced into economics by Spence (1973) to explain how seemingly self-defeating behaviors (e.g., knowingly pursuing a degree in a field one is unlikely to use) can still benefit rational agents by signaling information about one’s preferences or abilities (e.g, that one is smart enough to graduate college) to others in a way that cannot be faked. Signaling was introduced to biology around the same time by Zahavi (1975) to explain an analogous class of animal traits and behaviors: phenotypes that seem to reduce fitness, such as peacocks growing elaborate tails or gazelles stotting.1 As the theory points out, these acts credibly communicate information to potential mates or predators precisely because of their self-handicapping effect; voluntarily wasting resources can credibly signal a wealth of resources to begin with. The central insight of signaling theory is that rational agents only take actions they expect to be beneficial on net. A behavior can constitute definitive proof that the actor expects its benefits (including revealing information to you, the potential observer) to outweigh its costs. This pinpoints why signaling can only be accomplished by behaviors that incur real net costs when faked. The apparent downside of signaling behaviors are precisely what establish their credibility: the cost structure ensures that deceptive types cannot send the signal with impunity. This self-policing logic is what enables signaling equilibria to extend credible communication to low-trust environments, where honesty can not be enforced (by, e.g, reputation). Figure 1 illustrates two focal cases of signaling equilibria. In the first case (subfigure a), two types of agents receive the same benefit from engaging in a signaling behavior, but incur different costs. From the perspective of an external observer, witnessing the behavior constitutes definitive proof that the agent is of the type with lower cost.

Discussion. Mental proofs are most important in contexts where these supervening mechanisms are not present, for example because institutions are, or have become, weak or misaligned. Students for countries with failed credentialing systems, for example, will be asked to provide more mental proofs compared to others. Potential romantic partners may demand more mental proofs from each other in cultures that can not, or do not, enforce standards of care.

Mental proofs are a powerful tool for establishing both individual relationships and the common knowledge necessary for effective group action. Conversely, the weakening of mental proof can significantly stunt people’s ability to form new interpersonal relationships and cooperative initiatives. This suggests that the vibrancy of mental proof is a precursor to the general health of social, familial, political, and economic life in a community. As the example of sincere apology makes clear, mental proof also play an important role in repairing and maintaining close ties. The general erosion of mental proof therefore also has the effect of diminishing the psychological and social benefits such relationships confer. Feeling that one is understood and cared about by another is a deep psychological priority (Cahn, 1990; Oishi et al., 2010; Reis et al., 2000; Lun et al., 2008), separate from the material benefits that care might bring. Our discussion of social proof also showed how mental proofs help people establish collective interests, beliefs, and capacities. These proofs provide more than just the psychological benefits of knowing one is not alone. Collective action matters—the creation and maintenance of groups is a basic feature of civil society, in general, and the success of democratic government, in particular (Putnam, 1994).

Artificial intelligence’s deleterious effects on both costly signaling and proofs of knowledge can be prevented if people can clearly delineate between communicative acts undertaken with and without the assistance of such tools. In the United States, this is a stated goal of a 2023 Executive Order (No. 14110, Section 4.5),6 and Jain et al. (2023) detail a variety of strategies to establish and maintain such distinctions, which they label the contextual confidence of communication. Clarifying the difference between AI and human content would enable us to reap the benefits of automation (where advantageous) while preserving the capacity of humans to harness the benefits of mental proof (where important).

Our analysis also suggests that the economic benefits, in terms of reduced labor costs, of automating human mental effort can backfire. This may be especially true for jobs where care and understanding are paramount. Consider a patient in talk therapy for a difficult-to-treat condition. Generative artificial intelligence may help the therapist by spotting patterns and dynamically recommending better strategies of engagement, but we do not yet understand how this might undermine the therapeutic alliance (Tal et al., 2023; Zetzel, 1956). Clarifying the work being done by mental proof may help guide efforts to surgically target aspects of these jobs that can benefit from automation without damaging their core efficacy. Our work has focused on the beneficial effects of mental proof. As Spence (1973) pointed out when introducing the concept, however, the existence of signaling equilibria can be socially costly and may create disadvantages for some participants. This suggests, in turn, that in some cases—which ones remains a subject for further research—outcomes may be improved when Generative AI destroys an equilibrium previously supported by mental proof.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Does AI-assisted work increase total productivity or just shift time? Can readers reliably distinguish AI-written text from human writing? Why do confident AI outputs mislead human trust calibration? Does AI deployment reduce or exacerbate workplace inequality and income instability? Can humans reliably detect and resist AI-generated misinformation? How do AI hiring systems affect authenticity, fairness, and candidate preferences? How do philosophical assumptions about AI consciousness affect practical harms and design? What social dynamics enable or prevent agent collusion? How do agents learn to distinguish valuable feedback from noise? Should governance of agentic AI systems be runtime or design-time? Why do LLM research ideation systems generate novelty but lack diversity? Can AI research automation sustain progress through accelerating feedback loops? How does tokenization reshape what we value in intelligence?