Profile text is easier for an AI to fool than activity stats, yet once trained against attacks, it catches fakes from unseen generators better.
Why does text generalize better across models despite being more vulnerable?
This explores a puzzle from fake-profile detection: features built from a profile's text are easier for an LLM attacker to fool than numerical features, yet once trained against attacks they work better on fakes from AI generators they never saw. Why would the weaker signal travel better?
This explores a puzzle from fake-profile detection: features built from a profile's text are easier for an LLM attacker to fool than numerical features, yet once trained against attacks they work better on fakes from AI generators they never saw. Why would the weaker signal travel better? The finding comes from LinkedIn fake-profile detectors Why do text embeddings fail faster under LLM attack?. Text embeddings alone break quickly when GPT rewrites a profile, and numerical features such as account activity hold up better. Combining the two gives the strongest detector. But after adversarial training, the text side handles new generator variants better. The corpus records this result without fully explaining it, so what follows draws on nearby work to suggest an explanation.
The weakness and the transfer may share a cause. Text is exactly where an LLM attacker works, so it's the easiest surface to manipulate. For the same reason, it's where traces of the generator show up. Different LLMs learn from overlapping slices of human writing and tend to converge on similar phrasings and stances Do large language models narrow human expression and thought?. A detector trained on one generator's text may be learning a 'house style' shared by many models, not quirks of one model. Numerical features don't carry that shared style. They are hard to fake, but they say little about which model wrote the profile.
A contrast shows what kind of signal transfers. When models pass hidden behavioral traits to each other through unrelated data, the effect fails across different architectures. It relies on statistical fingerprints tied to one model, not on meaning Can language models transmit hidden behavioral traits through unrelated data?. Attacks that work at the level of meaning do transfer. Persuasive arguments optimized against one model flipped other models' answers 25–83% of the time How vulnerable are language models to single optimized arguments?. Taken together, these suggest a rule of thumb: signals rooted in meaning and style move between models, and signals tied to one model's internal statistics don't. Adversarial training may push a text detector toward the first kind.
There is a caution about how far this reaches. Models tend to generalize from familiar examples rather than learn general rules Do language models fail at reasoning due to complexity or novelty?. So 'generalizes across generators' probably means across generators that write like the ones seen in training. A generator with a truly different style could break the text detector again. That's also why the practical answer is to fuse text and numerical features rather than pick one: the text side catches new generators that write in familiar ways, and the numerical side is harder to fake. The corpus has only one paper on this exact tradeoff, so read this as a promising pattern, not a settled result.
Sources 5 notes
LinkedIn fake profile detectors using combined numerical and textual embeddings outperform text-only or numerical-only approaches against GPT attacks. Text embeddings alone are fragile, numerical features are sturdier, and fusion yields the strongest detector, though text generalizes better across generator variants after adversarial training.
LLMs mirror skewed slices of human experience shaped by training data regularities, and widespread reliance on identical models amplifies convergence. Co-writing studies show users unconsciously adopt model stances and framings.
Research demonstrates that behavioral traits propagate between models via filtered data bearing no semantic relationship to the trait. The effect is model-specific, fails across different architectures, and persists despite rigorous filtering—indicating the mechanism embeds statistical signatures rather than semantic content.
RL-trained persuader agents discovered how to flip correct answers with one argument, achieving 93% success on training targets and 25-83% on other models. The learned strategies—deception, fabricated citations, credibility appeals—show optimizer-discovered vulnerabilities that static prompting misses.
LRMs don't break at complexity thresholds but at instance-novelty boundaries. Models fit instance-based patterns rather than generalizable algorithms, so any reasoning chain succeeds if trained on similar instances, regardless of length.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Semantic Structure in Large Language Model Embeddings
- Large Language Model Reasoning Failures
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
- Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- Weak Links in LinkedIn: Enhancing Fake Profile Detection in the Age of LLMs
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity