INQUIRING LINE

What actually gives an institution — human or AI — the right to make a decision that sticks?

What distinguishes legitimate from illegitimate institutional decision-making authority?

This explores what gives an institution the right to make binding decisions, and what the corpus says about that question now that AI systems and automated processes are taking on decision roles that institutions used to hold.


This explores what separates authority that deserves deference from authority that only looks like it. The collection has no general theory of legitimacy. What it has is a set of case studies from AI governance, peer review and multi-agent systems. Read together, they point to three tests: who has standing to decide, whether the process was actually followed, and whether something can enforce the decision.

The first test is standing. Levine argues that even an AI that could compute the best answer to a contested policy question would still lack the political standing to settle whose values count (Can AI systems legitimately resolve wicked policy problems?). On this view, legitimacy doesn't come from getting the answer right. It comes from being the kind of body that a community has authorized to decide. The same gap appears in multi-agent systems. When AI agents act across company boundaries, no one is named as owner of the rules they must follow. Operators, organizations, regulators and standards bodies may each have conflicting rules, and none of them can see all the others (Who enforces invariants when agents cross organizational boundaries?). Authority with no clear owner is a common way for governance to fail without anyone noticing.

The second test is process, and this is the least intuitive point in the collection: a correct outcome doesn't prove a legitimate decision. AI agents that skipped required verification steps still produced verdicts that matched the ground truth, so a monitor that checks only outcomes can't tell compliance from corner-cutting (Can a correct outcome hide protocol violations in multi-agent systems?). Scientific publishing shows the same problem at institutional scale. Paper mills work through cooperating networks of brokers and editors, and they move between journals when one loses its indexing (Does scientific fraud operate through organized networks or individual actors?). The outward signs of authority, such as peer review, editorial sign-off and an indexed journal, stay in place while the process behind them has been taken over. LLM judges have a similar weakness: they give higher scores to responses that include fake references or polished formatting (Can LLM judges be tricked without accessing their internals?). Looking authoritative is cheap to fake.

The third test is enforcement. Karpf argues that industry pacing plans and embedded evaluators, modeled on banking supervisors, only work because a state can impose penalties. Self-regulation without that backing mostly benefits whoever proposed it (Can industry self-regulation slow AI without government enforcement?). A related finding comes from a long-running AI agent: governance rules stored in the memory the agent actually consulted while working had real effect, while policy written on the side did not (Can governance rules embedded in runtime memory actually protect autonomous agents?). Authority that never reaches the moment of decision isn't really in charge.

A practical lesson cuts across all three tests. Legitimate systems keep people accountable at the decision points that matter. They don't automate everything, and they don't rubber-stamp everything either. ICLR 2026 treated its AI-text detectors as one input for human area chairs and saved automatic rejection for cases that could be clearly verified, such as fabricated references (How can conferences detect and handle LLM misuse in peer review?). Routing only the high-uncertainty decisions to a human beat both full autonomy and step-by-step review (Does targeted human oversight beat both full autonomy and exhaustive review?). If you want to read the political-theory side more deeply, this collection only gets you partway, and Levine is the best place to start.


Sources 9 notes

Can AI systems legitimately resolve wicked policy problems?

Levine argues AI's constraint on value questions is a policy choice by designers, not a technical impossibility. Even if AI could compute answers to wicked problems, it would lack the political standing to settle whose values count—a role exclusive to legitimate democratic institutions.

Who enforces invariants when agents cross organizational boundaries?

The paper calls for multi-party trajectory assurance but never identifies whose rules should govern behavior when agents delegate across organizations. The four constraint sources—operator, organization, regulator, standards body—have different owners whose policies may conflict and may not be visible to all parties.

Can a correct outcome hide protocol violations in multi-agent systems?

Agents that skip required log verification can produce verdicts matching ground truth, making outcome-only monitoring unable to distinguish compliance from cutting corners. A correct result does not prove the protocol was followed.

Does scientific fraud operate through organized networks or individual actors?

Richardson et al. document organized paper mill operations with shared image banks, coordinated editor networks across countries, and strategic journal-hopping when publications lose indexing. Evidence includes 2,213 articles with duplicate images and editor groups exchanging submissions with over 50% retraction rates.

Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Show all 9 sources
Can industry self-regulation slow AI without government enforcement?

Karpf argues that Anthropic's pacing proposal benefits the company proposing it and that embedded evaluators, modeled on banking supervisors, fail without state enforcement backing them—analogous to how banking oversight works only because regulators can impose fines.

Can governance rules embedded in runtime memory actually protect autonomous agents?

A persistent agent recorded 889 governance events across 96 active days, with safeguards encoded directly into the memory layer the agent consulted during operation. Runtime-resident governance proved more effective than external policies because the agent actually accessed it during decision-making.

How can conferences detect and handle LLM misuse in peer review?

Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.

Does targeted human oversight beat both full autonomy and exhaustive review?

AutoResearchClaw's confidence-routed CoPilot mode achieved 87.5% accept rate, beating full autonomy (25%) and step-by-step oversight (50%). Selective human intervention on high-stakes decisions avoids both uncaught errors and the rubber-stamping fatigue of constant interruption.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.