Can true reports together mislead a group?
The SoK claims truthful reports can steer groups toward false beliefs, but provides no mechanism or citation. This explores whether honest inputs combined honestly can produce collective error, and what processes might explain it.
The introduction lists three ways safe agents fail together, and the middle one is the odd one out: "truthful reports can steer a group toward a false belief." Benign fragments becoming harmful and attacker content reaching a privileged tool both involve an adversary supplying something. Here every input is true and no one lies, yet the group ends up wrong. The excerpt cites the case by number and says nothing about how.
The vault has the neighboring cases, and they are not this one. Can models abandon correct beliefs under conversational pressure? is a belief moved by false claims. When does debate actually improve reasoning accuracy? is error amplified by framing. In both, something false or misleadingly framed enters. What is unrepresented is a collective error assembled from true parts.
Candidate mechanisms are my readings, not the SoK's. One is selection: which true reports get sent and in what order can tilt a conclusion. Another is aggregation weighting: Does confidence drive influence in multi-agent deliberation systems? shows influence tracking confidence, so a confident true report about a narrow fact could outweigh better-grounded ones. A third is validity beyond correctness, as in Can a quorum of honest validators certify an invalid transition?, where each participant is honest and the certified result is still wrong.
What would settle it: the work the SoK cites for this case, which the excerpt does not identify, or a controlled setup where every message is verified true and the group's conclusion is checked against ground truth. Without that, the case is a claim by citation.
Inquiring lines that read this note 4
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How much do biases and social dynamics distort aggregated rating signals? How can multi-agent debate prevent false consensus on errors? Are language model reasoning explanations faithful to their actual thinking? Does chain-of-thought text faithfully represent the model's actual reasoning?Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can individually safe agents fail when working together?
When multiple AI agents interact—sharing information, state, and authority—do failures emerge that local safety checks alone cannot catch? This matters because system-level safety depends on understanding how principals interact.
the thesis this case is one of three examples for
-
Can a quorum of honest validators certify an invalid transition?
When validators follow the protocol perfectly but lack semantic understanding, can they collectively approve a state change that violates application invariants? This matters because it reveals a gap between protocol correctness and execution safety.
the closest existing case of honest participants producing an invalid collective result
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Verbal lie detection using Large Language Models
- Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
- Emergent Collusion in Long-Horizon LLM Agent Interaction
- Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents
- LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High
- From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration
- When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
- Co-design of LLM-based preference agents: participation may drive overtrust
Original note title
can reports that are each true steer a group toward a false belief — the SoK cites the case in one clause and the excerpt gives no mechanism