If AI takes over the jobs that quietly kept companies answerable to people, who's still watching the store?
Does the loss of human labor participation in systems undermine alignment?
This explores whether AI systems and the institutions around them stay pointed at human interests partly because they need people to do the work, and what happens to that safeguard when the people are replaced.
This explores whether human jobs inside a system quietly do alignment work, so that replacing them with AI removes a check nobody designed on purpose. The clearest case in the corpus is the 'gradual disempowerment' argument Does incremental AI replacement erode human influence over society?. Economies, governments and cultural institutions stay roughly aligned with what people want partly because they depend on workers who care how things turn out. That dependence is implicit alignment: a firm or agency can only drift so far from human preferences before the people running it push back or leave. As AI takes over that labor, the check weakens without anyone deciding to remove it. Because institutions depend on each other, their drift can compound until it becomes hard to reverse. The replacement is also uneven. Firms with more exposure to AI swap out online freelance workers faster and more cheaply than others do Do firms substitute labor for AI at different rates?. So the erosion probably shows up first in pockets, not evenly across the economy.
The same pattern appears at the level of individual models, under different names. A review of more than 400 alignment papers found that the field almost always treats alignment as changing the AI's behavior. Very little work looks at how humans adapt to AI, and the review warns that this gap slowly wears down people's ability to oversee these systems Why does alignment research ignore how humans adapt to AI?. Jan Leike makes a related point from inside a lab. Today's alignment successes are 'easy mode' because humans can still read and audit what models do. Once models act in ways people can't follow, the hard problem comes back Can we solve AI alignment before models become uninterpretable?. Put these together and labor and oversight look like versions of the same thing: humans who stay close enough to the work to notice when it goes wrong.
A more philosophical thread explains why this closeness might matter. One argument, drawing on the philosopher Charles Peirce's theory of signs, holds that goals written only in symbols can't be guaranteed to match real values. They need contact with the world and with other people Can AI systems achieve real alignment without world contact?. Human workers are one of the main ways AI systems get that contact. Work on self-improvement reaches a similar conclusion from another direction: reliable improvement needs checks from outside the system, because a model can't fully verify itself What actually constrains AI systems from learning misalignment?. Another argument says AI should follow the norms of particular social roles, such as doctor or teacher, instead of averaged-out preferences Should AI alignment target preferences or social role norms?. That raises an open question the corpus doesn't settle: who keeps a role's standards alive once people no longer fill it?
One note complicates this. Simply having humans take part doesn't produce alignment. In a 12-person study, people who helped design their own AI preference agents felt well represented. Independent testing found the agents were only partly aligned, and more generic than the people's own answers. Taking part created the feeling of alignment without the substance Does co-design participation hide misalignment in preference agents?. So keeping humans involved isn't enough on its own terms. What protects alignment is humans whose involvement can actually change outcomes and catch mistakes. Losing labor matters because it usually removes that, but a human in the loop as a token gesture doesn't bring it back.
A caveat: only one note in the collection takes on the labor question directly. The rest connect to it from oversight, grounding and verification. The idea that work itself functions as an alignment mechanism is still an open claim here, not an established result.
Sources 8 notes
Societal systems stay aligned partly through dependence on human workers who care about outcomes. As AI replaces this labor, explicit alignment controls weaken and systems drift from human preferences. Interdependent misalignment across institutions could become irreversible.
Higher AI-exposed firms replace online labor marketplace workers with AI tools faster and at lower cost than less-exposed firms, suggesting returns to scale in internal AI capability rather than uniform technology diffusion.
A 400+ paper review shows alignment overwhelmingly targets AI behavior change while human-to-AI adaptation receives minimal attention. This creates vulnerabilities like specification gaming and erodes human capacity for oversight over time.
Leike reports that simple interventions reduced agentic misalignment to near zero in recent models through automated auditing metrics, but this success depends on human interpretability; once models act in ways humans cannot understand, alignment becomes an unsolved hard problem.
Peircean semiotics reveals that symbolic goal encoding without world contact and social mediation cannot guarantee correspondence to actual values. LLMs operating in pure symbol manipulation risk divergence between stated goals and real-world outcomes.
Show all 8 sources
Alignment philosophy is shifting from matching human preferences to enforcing role-appropriate standards. Self-improvement remains bounded by the generation-verification gap, meaning reliable improvements require external oversight rather than learned metacognition.
Preferentialist alignment approaches fail because preferences don't capture thick moral values, uniform aggregation produces epistemic injustice, and preference optimization creates systematic misalignment with social roles. Contractualist alignment negotiated by stakeholders and bounded by supra-national, organizational, and individual levels works better.
In a 12-person study, participants felt their co-designed preference agents represented them well, but independent validation revealed mixed alignment and agents that were more generic and abstract than human responses. The co-design process itself—through transparency, limited testing, and cognitive biases—appears to have produced the feeling of alignment rather than ensuring actual alignment.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Position: Towards Bidirectional Human-AI Alignment
- Beyond Preferences in AI Alignment
- An Alien Mind
- Alignment is not solved but it increasingly looks solvable
- Position: Anthropomorphic Misalignment Research Needs Stronger Evidence
- Sycophancy Towards Researchers Drives Performative Misalignment
- Stress Testing Deliberative Alignment for Anti-Scheming Training
- Conversational Alignment with Artificial Intelligence in Context