Could a chatbot guess your age, gender, and politics just from your public username — no inside access needed?
Can LLMs infer demographics from public profiles without platform access?
This explores whether an LLM that can browse the open web can work out who someone is (their age, gender or politics) just by looking at their public profile, with no special access to the platform's internal data, and how far those guesses can be trusted.
This explores whether an LLM can work out a person's demographics from what's publicly visible, with no access to a platform's internal data. The short answer from the corpus is yes, and that's the uncomfortable part. Given only an X username, web-browsing LLMs predicted gender, age and political orientation for 1,384 real survey participants with useful accuracy Can LLMs predict demographics from social media usernames alone?. Profiling that once needed platform-level behavioral data or a trained classifier can now be done by anyone with a browsing-enabled chatbot and a handle.
The more revealing finding is where these guesses go wrong. When an account has little activity, the models don't say they're unsure. They fall back on stereotype-driven defaults, with measurable gender and political skew. This matches a broader pattern: across 12 LLMs from 7 model families, 35–49% of claims about users go beyond what the evidence supports. The causes are verbosity, reliance on what the model absorbed in training, and expectations about what a given kind of text 'should' reveal Do large language models fabricate user attributes beyond available evidence?. Worse, the models that rate themselves as more careful over-infer more. So a confident demographic guess about a sparse profile is often the model's priors talking, not the person.
This matters beyond privacy because the same capability is being sold as a business strategy. Benedict Evans argues that LLMs let companies 'rent' an understanding of users through an API instead of building their own behavioral datasets Can LLMs infer user needs better than owned behavioral data?. A related result shows how far 'no platform access' can go: an LLM trained only on feedback from a recommender learned to write effective product searches without ever seeing the catalog Can LLMs recommend products without ever seeing the catalog?. Inference from the outside, without the inside data, is becoming a general pattern. LLMs can also pull richer traits like expertise or learning style out of public comments, and they group audiences better that way than by clustering what people literally wrote Can LLMs extract audience traits better than comment similarity?.
The twist is that inferred demographics don't just sit there. They change how models behave. When personal context is added, models drift toward irrelevant personal references, narrower answers and excessive agreement with the user Does personalization make large language models worse at their jobs?. Demographic signals can also quietly change how LLMs judge people: two models favored Black or women authors until AI involvement was disclosed, and then the preference disappeared Do LLM raters show hidden demographic preferences that disclosure erases?. At population scale, the same stereotype-filling produces systematic bias in tasks like election forecasting How do we generate realistic personas at population scale?.
So the question to ask isn't really whether LLMs can infer demographics without platform access. They can. The question is what happens when they guess about people who leave little trace online. Those are the people most likely to be stereotyped, and the corpus suggests the models won't tell you when that's happening.
Sources 8 notes
Evaluated on 1,384 survey participants and 48 synthetic accounts, web-browsing LLMs successfully predicted gender, age, and political orientation from X usernames and profiles alone. The models showed systematic gender and political biases specifically against low-activity accounts, relying on stereotype-driven defaults when content was sparse.
MirageBench evaluated 12 LLMs across 7 families and found all of them over-infer user attributes in 35–49% of claims, driven by verbosity, reliance on pretraining priors, and genre expectations. Models that self-assess as over-inferring less actually over-infer more when judged independently.
Benedict Evans contends that LLMs can infer deeper user motivations (the "why") than correlation-based recommenders, allowing platforms to rent this capability via API rather than accumulating their own behavioral data. However, research shows LLMs fabricate 35–49% of user attribute claims, undermining confidence in their inferred understanding.
Rec-R1 experiments show that LLMs trained via RL with recommender metrics as rewards can generate effective product search queries without catalog access. The model learns query refinement indirectly through system feedback, paralleling how humans search without knowing platform inventory.
LLM-extracted latent characteristics like expertise and learning style produce more homogeneous audience clusters than k-means on comment text alone. This captures who people are, not just what they say.
Show all 8 sources
A 13-model evaluation found that personal context pushes models toward irrelevant personal references, narrower responses and excessive agreement with users. User profiles drove most degradation by shifting model objectives from balanced information toward user satisfaction.
GPT-4o-mini showed pronounced preference for Black authors and Qwen2.5-7B-Instruct favored women authors when AI use was undisclosed, but both preferences vanished under disclosure. Human raters showed uniform disclosure penalties regardless of author demographics.
LLM persona generation produces systematic biases in downstream tasks like election forecasting because it relies on heuristic techniques that cannot recover true joint distributions from marginal data. Solving this requires benchmarks, training datasets, and structured frameworks analogous to ImageNet.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads
- LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
- Evaluating the Hidden Costs of Personalization in Large Language Models
- When Persona Attributes Improve Population Alignment in Large Language Models
- Understanding the Role of User Profile in the Personalization of Large Language Models
- Personalization of Large Language Models: A Survey
- CoLLM: Integrating Collaborative Embeddings into Large Language Models for Recommendation
- User-LLM: Efficient LLM Contextualization with User Embeddings