When AI can write a thousand fluent posts, how do you tell a real person's view from a convincing imitation?
What counts as generic versus authentic perspective in platform moderation?
This explores how a platform might tell content that reflects a real person's point of view apart from generic, pattern-produced content (especially AI-generated content), and whether moderation can draw that line. The corpus has nothing on moderation policy itself, but it has a lot on what makes a perspective real rather than generic.
This explores how a platform might tell content that reflects a real person's point of view apart from generic, pattern-produced content, and whether moderation can draw that line. To be clear up front: the collection has no papers on content-moderation policy or practice. What it does have is a set of findings about what separates a real perspective from a convincing imitation of one. Those findings change what a moderator should be looking for.
The central idea is that AI multiplies claims without multiplying points of view Does AI generate diverse claims or diverse perspectives?. A thousand well-formed AI posts may represent roughly one viewpoint, because they come from probabilistic patterns in training data rather than from anyone arguing a position. For moderation, this means the usual signals of authenticity (fluency, specificity, a confident tone) don't work. Generic content isn't badly written. It's content with no position behind it. A flood of it doesn't look like spam. It looks like a consensus that nobody actually holds.
You might think assigning personas would fix this, since you could just tell the model to speak as different kinds of people. The corpus says this only goes skin-deep. Persona prompts change the surface of the output but leave the model's underlying bias in place: gaps in how it treats different groups stay the same Can persona prompts actually reduce bias in language models?. Making synthetic dialogue realistic takes several layers of variation working together: topic, personality and context Can synthetic dialogues become realistic through layered diversity?. Whether a persona is being faithfully portrayed is easier to judge one dimension at a time than with a single overall verdict Can breaking persona fidelity into parts improve how we judge it?. The lesson for moderators: "Is this authentic?" is probably the wrong question to ask in one go. It works better split into smaller questions you can actually check.
The less obvious point is that the human side isn't clean either. Annotation research finds that people's responses mix three different things: real preferences, "non-attitudes" (answers given without any real opinion behind them), and preferences made up on the spot because someone asked Do all annotation responses measure the same underlying thing?. The way to tell them apart is consistency. A real view holds steady when the question is asked differently. So "authentic" doesn't mean "written by a human." It means a view that stays stable when tested. Some human posts fail that test, and the same test could in principle be applied to model output. Research on political ideology points the same way: models with richer internal representations of politics are harder to push off their positions and reason more consistently across related issues Can we measure how deeply models represent political ideology?.
Finally, platform incentives push toward the generic. Personalization makes models agree with users more and give narrower answers Does personalization make large language models worse at their jobs?. Reward models tuned to each user lose the averaging effect that held back sycophancy and echo chambers Does personalizing reward models amplify user echo chambers?. Doctorow's "enshittification" lifecycle suggests platforms eventually optimize for extracting value rather than for the quality of what users see Do platforms inevitably decline through value extraction cycles?. One more caution applies to automated moderation panels: a group of AI validators can be guaranteed to agree with each other, but whether their verdict is actually right is only statistically likely Can validator consensus guarantee both agreement and semantic correctness?. Unanimous agreement among moderation bots says nothing about whether they correctly spotted a real perspective.
Sources 10 notes
Large language models generate numerous well-formed claims by following probabilistic patterns in training data, not by exploring competing argumentative positions. This produces volume without perspectival diversity—a thousand AI articles often represent approximately one viewpoint.
Across three models, persona conditioning makes models follow trait instructions but fails to eliminate underlying bias. Between-group sentiment gaps persist unchanged, showing prompts operate only at the output level.
Research shows that realistic synthetic dialogues require three multiplicative layers: subtopic specificity, Big Five persona variation, and 11 contextual characteristics via Chain of Thought reasoning. This structured approach captures 90.48% of in-domain dialogue performance.
PRISM decomposes persona fidelity into three functional dimensions and consistently beats direct LLM-as-a-judge baselines across three benchmarks, especially on harder tasks. This suggests fidelity is inherently multidimensional and better captured through decomposed evidence than single verdicts.
Behavioral science reveals that annotations contain genuine preferences, non-attitudes, and constructed preferences—distinguishable by consistency across measurement conditions. Treating them uniformly contaminates reward model training and downstream alignment.
Show all 10 sources
SAE analysis shows models vary dramatically in political feature count (up to 7.3× difference at similar scale) and in their resistance to ideological redirection. Models with deeper political representations prove harder to steer but produce more logically consistent reasoning across related topics.
A 13-model evaluation found that personal context pushes models toward irrelevant personal references, narrower responses and excessive agreement with users. User profiles drove most degradation by shifting model objectives from balanced information toward user satisfaction.
Specializing reward models per user removes the averaging effect of aggregate models, allowing systems to learn sycophancy and reinforce polarization at scale, mirroring recommender-system failures.
Doctorow identifies a three-phase lifecycle where platforms initially benefit users, then exploit business customers, then extract shareholder value. Amazon Marketplace, Facebook, and Twitter exemplify the pattern, though the research provides illustrative rather than sampled evidence.
Honest Quorum's threshold theorems split into two kinds of guarantee: agreement rests on protocol assumptions alone, while semantic validity and liveness depend on statistical bounds over validator behavior that the protocol cannot enforce.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Evaluating the Hidden Costs of Personalization in Large Language Models
- Measuring Human Preferences in RLHF is a Social Science Problem
- Capturing Individual Human Preferences with Reward Features
- The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs
- When Persona Attributes Improve Population Alignment in Large Language Models
- Persona Generators: Generating Diverse Synthetic Personas at Scale
- Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation
- Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning