Can fake profile detectors catch GPT-generated LinkedIn profiles?
Text-based detectors perform well on manually created fakes but struggle with GPT-generated ones. This matters because attackers can now use LLMs to create convincing fake profiles at scale.
The paper's central claim is that text-based fake profile detectors built on genuine and manually created fakes do not survive LLM-generated fakes. The abstract says the existing detectors are "highly effective in detecting manually created fake profiles (False Accept Rate: 6 −7%)" but "fail to identify GPT-generated profiles (False Accept Rate: 42−52%)". The introduction gives the prior SSTE method of Ayoobi et al. about F1 96% against manual fakes and about F1 68% against LLM-generated ones. To test this, the authors generated 600 fake profiles with GPT-4-Turbo. The Discussion attributes the degradation to high textual similarity between GPT-generated and real profiles (mean 88.9%).
The countermeasure is GPT-assisted adversarial training, in which GPT-generated profiles are added to the training data. The abstract reports that this restores the false accept rate to between 1 and 7 percent "without impacting the False Reject Rates (0.5 −2%)". The Discussion separates the sources: GPT-3.5-assisted training cut false accepts on GPT-3.5 profiles to as low as 1.3% but left the detector exposed to GPT-4 profiles (16.9% to 19.3%), and the reverse held for GPT-4-assisted training. Training on the combined GPT-3.5 and GPT-4 data gave false accept rates of 1.34% to 2.6% with F1 consistently above 97.5%. Flair embeddings with XGBoost scored best, at F1 98.2% and false accept rate 1.34% on combined attacks. The excerpt reports these results but does not explain why the mixed training set generalizes where single-generator sets do not.
Against the nearest notes, this is a detector failing on the wrong distribution, and it fails in the opposite direction from Why do fake news detectors flag AI-generated truthful content?. That note reports detectors flagging LLM-written text as fake. Here the GPT-generated fakes pass as legitimate. The excerpt's human and GPT-4 comparison also parallels Can LLM judges be fooled by fake credentials and formatting? in one respect: an LLM used as the judge is a weak detector on its own. GPT-4 reached F1 71.3% zero-shot and 85.7% few-shot, and human annotators reached F1 58.9%, all below the adversarially trained models.
The excerpt does not establish robustness beyond its test setting. Its own limitations section lists English-only LinkedIn profiles, no test on other platforms, a fixed set of encoders, and attacks and adversarial training that both used models from the OpenAI GPT family. Legitimate profiles that people wrote with LLM help were not tested. The implication is narrow. The 1 to 7 percent false accept rates describe robustness to GPT-family fakes on this data, not to LLM-generated fakes in general. A platform adopting this training would need evidence against other model families before treating the restored rates as a general fix.
Inquiring lines that read this note 10
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How does AI-generated content create social proof without authentic interaction?- How often does LinkedIn wrongly flag legitimate posts as AI-generated?
- How similar are GPT-generated fake profiles to real human profiles?
- How does LinkedIn's verification system affect what content appears in feeds?
- Can platforms trust these detection rates on profiles created with human-LLM collaboration?
- What triggers LinkedIn's detection of inauthentic content from heavy AI use?
- How does LinkedIn's platform response address detected AI-generated content?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do fake news detectors flag AI-generated truthful content?
Fake news detectors may systematically misclassify LLM-generated text as deceptive. We explore whether this bias stems from detecting AI style rather than actual falsehood, and what that means for detection accuracy.
opposite error direction: that detector over-flags LLM text as fake; this one passes LLM-generated fakes as real
-
Can LLM judges be fooled by fake credentials and formatting?
Explores whether language models evaluating text fall for authority signals and visual presentation unrelated to actual content quality, and whether these weaknesses can be exploited without deep model knowledge.
both test an LLM as the evaluator; here GPT-4 is a weaker detector than trained models
-
Why do text embeddings fail faster under LLM attack?
Text-only profile embeddings collapse under LLM adversarial attacks while numerical features remain stable. Understanding this gap could reveal whether the fragility stems from how text encodes meaning or from something specific to how LLMs generate profiles.
sibling note: shows which feature types the adversarial training fix works through
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Weak Links in LinkedIn: Enhancing Fake Profile Detection in the Age of LLMs
- Mapping the Increasing Use of LLMs in Scientific Papers
- Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media
- AI Now Writes as Many Online Articles as Humans
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- GPT-fabricated scientific papers on Google Scholar: Key features, spread, and implications for preempting evidence manipulation
- Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
- Keeping conversations real on LinkedIn
Original note title
LinkedIn fake profile detectors miss GPT-generated profiles at 42 to 52 percent false accept rates and GPT-assisted adversarial training restores them