INQUIRING LINE

If detectors for fake LinkedIn profiles are retrained on one AI's fakes, do they still catch fakes from other models?

How would text and numerical features perform against non-GPT profile generators?

This explores whether fake-profile detectors that combine text features (what a profile says) with numerical features (counts, dates and other structured fields) would still catch fakes written by language models other than GPT, given that the research tested them mainly against GPT.


This explores whether the detection lessons learned from GPT-written fake LinkedIn profiles would carry over to fakes made by other language models. The short answer is that the corpus doesn't test this directly. Its evidence comes from GPT-based attacks. Still, it has useful clues about which kinds of features are likely to hold up when the generator changes.

Start with the baseline. Detectors trained on real profiles and on fakes written by hand miss GPT-generated profiles 42–52% of the time. Once they are retrained on GPT fakes, that miss rate falls to 1–7%, and real users aren't wrongly flagged any more often than before Can fake profile detectors catch GPT-generated LinkedIn profiles?. So detectors are not doomed. They mostly fail on kinds of fakes they have never seen. That is exactly the worry with non-GPT generators: each new model family could be another kind of fake the detector hasn't seen.

The feature-level findings matter here. Text-only detectors break down fastest under LLM attack, because fluent generated text is what these models are best at faking. Numerical features, such as connection counts, account age and how complete the profile is, are sturdier, because a language model doesn't control them as easily. Combining the two gives the strongest detector overall Why do text embeddings fail faster under LLM attack?. The same research has a twist that bears directly on your question: after adversarial training (retraining on generated fakes), the text features generalize *better* than the numerical ones across different generator variants. Text is weak against a new attack, but once the detector learns what generated prose looks like, that knowledge transfers. This suggests that a combined detector retrained on several generators is the best bet against unfamiliar ones. Retraining only on GPT probably leaves gaps.

There is also a reason to expect some transfer across model families, though this is inference, not something the corpus tested. A study of 12 LLMs from 7 different families found they all share the same habits: they are wordy, they lean on what they absorbed in training, and they follow genre conventions Do large language models fabricate user attributes beyond available evidence?. If the profiles different models write carry a similar fingerprint, a text detector trained on one family may partly recognize the others. Making that training data may not need real examples of every attacker either. Related work shows that a short text description of a target domain can be enough to generate synthetic training data that adapts a retrieval model without any real target data Can you adapt retrieval models without accessing target data?. It's an analogous technique, not a proven fix for fake profiles.

The honest gap: no note in the collection measures detection rates against Llama, Claude, Mistral or other non-GPT generators. The useful takeaway is that in this research, a detector's weakness came less from which features it used and more from which generators it was trained on. Combined features plus retraining on many different generators is the strategy the evidence points toward.


Sources 4 notes

Can fake profile detectors catch GPT-generated LinkedIn profiles?

Detectors trained on genuine and manual fakes miss GPT-generated profiles at 42–52% false accept rates, but adversarial training on GPT-generated data restores detection to 1–7% false accepts without raising false rejects.

Why do text embeddings fail faster under LLM attack?

LinkedIn fake profile detectors using combined numerical and textual embeddings outperform text-only or numerical-only approaches against GPT attacks. Text embeddings alone are fragile, numerical features are sturdier, and fusion yields the strongest detector, though text generalizes better across generator variants after adversarial training.

Do large language models fabricate user attributes beyond available evidence?

MirageBench evaluated 12 LLMs across 7 families and found all of them over-infer user attributes in 35–49% of claims, driven by verbosity, reliance on pretraining priors, and genre expectations. Models that self-assess as over-inferring less actually over-infer more when judged independently.

Can you adapt retrieval models without accessing target data?

Research demonstrates that a brief textual domain description suffices to generate synthetic training data for retrieval fine-tuning, outperforming baselines in zero-target-access scenarios and enabling adaptation where conventional methods are blocked.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.