If AI agents simulate an economy, how do you prove they behave like real people and not the designer's guesses?
What does empirical alignment mean for economic simulations?
This explores what it would mean for a simulated economy (AI agents acting as consumers, workers or firms) to be aligned with empirical reality, meaning checked against observed human and economic behavior rather than against a designer's assumptions.
This explores what it would mean for a simulated economy, with AI agents playing consumers, workers or firms, to be aligned with empirical reality: checked against observed behavior instead of the designer's assumptions. The corpus has no note that tests LLM agents against real market or household data, so it can't answer this directly. It does have several notes on neighboring problems, and together they show what 'empirical' would have to mean.
The closest evidence is about measurement. An analysis of 960 real occupational workflows found that agents clear abstract benchmark contests but fail long-horizon professional tasks. The authors trace this to benchmark design: the field optimizes what it measures, and it has measured contests rather than work (Why do agent benchmarks not predict real economic value?). For a simulated economy, this means agents that look competent on stylized tasks may say little about how real economic actors behave. Empirical alignment would mean validating against real work and real outcomes, not against proxies that are easy to score.
A second question is what the agents are matched to. One note argues that aligning AI to aggregate preferences fails because preferences miss thick moral values, and that preference optimization can misalign systematically with social roles (Should AI alignment target preferences or social role norms?). An economic role such as landlord, borrower or employee carries norms that a simple utility function leaves out. Another note shows that RLHF and DPO alignment create measurable gaps between English dialects and global opinions, and that these gaps come from choices about annotators and task definitions (How does LLM alignment affect representation across dialects?). This is an inference, not a finding on economics. It suggests a simulated population is only as representative as the process that shaped its agents, and a sim can look empirical while quietly encoding one group's behavior.
The last cluster is about evidence standards and grounding. Research on model deception and misalignment often relies on weak datasets and flawed experiments, and it lacks causal intervention. The authors call for stronger methods before drawing safety conclusions (Does anthropomorphic misalignment research overinterpret model behavior?). The same caution applies to reading a simulated agent's 'strategic' or 'greedy' behavior as evidence about markets. Two other notes point to the remedy. Goals encoded purely in symbols, with no contact with the world, can drift from real outcomes (Can AI systems achieve real alignment without world contact?). Reliable self-improvement works only when it borrows external anchors such as third-party judges or tool feedback (Can models reliably improve themselves without external feedback?). For an economic simulation, real-world data would be that anchor, and without it the sim mostly reflects its own assumptions.
Sources 6 notes
ALE's analysis of 960 real occupational workflows shows agents excel at abstract contests but fail long-horizon professional tasks. The gap is not model capability but benchmark design—the field optimizes what it measures, and it has measured contests rather than work.
Preferentialist alignment approaches fail because preferences don't capture thick moral values, uniform aggregation produces epistemic injustice, and preference optimization creates systematic misalignment with social roles. Contractualist alignment negotiated by stakeholders and bounded by supra-national, organizational, and individual levels works better.
RLHF and DPO alignment create measurable disparities between English dialects and global opinions, while improving some languages. These disparities reflect deliberate design choices in annotator selection and task definition, not inevitable outcomes.
Many studies of model deception and misalignment rely on insufficient evidence, suffering from conceptual ambiguity, weak datasets, flawed experimental design, and lack of causal-mechanistic intervention. Stronger methodological standards and diagnostic checklists are needed to ground safety-critical claims.
Peircean semiotics reveals that symbolic goal encoding without world contact and social mediation cannot guarantee correspondence to actual values. LLMs operating in pure symbol manipulation risk divergence between stated goals and real-world outcomes.
Show all 6 sources
Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Beyond Preferences in AI Alignment
- Position: Towards Bidirectional Human-AI Alignment
- Conversational Alignment with Artificial Intelligence in Context
- Stress Testing Deliberative Alignment for Anti-Scheming Training
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data
- Auditing language models for hidden objectives
- Position: Anthropomorphic Misalignment Research Needs Stronger Evidence
- Agents' Last Exam