SYNTHESIS NOTE
Topics›Expertise in the Age of AI Content›this note

Can fake profile detectors catch GPT-generated LinkedIn profiles?

Text-based detectors perform well on manually created fakes but struggle with GPT-generated ones. This matters because attackers can now use LLMs to create convincing fake profiles at scale.

Synthesis note · 2026-10-06 · sourced from Expertise in the Age of AI Content

The paper's central claim is that text-based fake profile detectors built on genuine and manually created fakes do not survive LLM-generated fakes. The abstract says the existing detectors are "highly effective in detecting manually created fake profiles (False Accept Rate: 6 −7%)" but "fail to identify GPT-generated profiles (False Accept Rate: 42−52%)". The introduction gives the prior SSTE method of Ayoobi et al. about F1 96% against manual fakes and about F1 68% against LLM-generated ones. To test this, the authors generated 600 fake profiles with GPT-4-Turbo. The Discussion attributes the degradation to high textual similarity between GPT-generated and real profiles (mean 88.9%).

The countermeasure is GPT-assisted adversarial training, in which GPT-generated profiles are added to the training data. The abstract reports that this restores the false accept rate to between 1 and 7 percent "without impacting the False Reject Rates (0.5 −2%)". The Discussion separates the sources: GPT-3.5-assisted training cut false accepts on GPT-3.5 profiles to as low as 1.3% but left the detector exposed to GPT-4 profiles (16.9% to 19.3%), and the reverse held for GPT-4-assisted training. Training on the combined GPT-3.5 and GPT-4 data gave false accept rates of 1.34% to 2.6% with F1 consistently above 97.5%. Flair embeddings with XGBoost scored best, at F1 98.2% and false accept rate 1.34% on combined attacks. The excerpt reports these results but does not explain why the mixed training set generalizes where single-generator sets do not.

Against the nearest notes, this is a detector failing on the wrong distribution, and it fails in the opposite direction from Why do fake news detectors flag AI-generated truthful content?. That note reports detectors flagging LLM-written text as fake. Here the GPT-generated fakes pass as legitimate. The excerpt's human and GPT-4 comparison also parallels Can LLM judges be fooled by fake credentials and formatting? in one respect: an LLM used as the judge is a weak detector on its own. GPT-4 reached F1 71.3% zero-shot and 85.7% few-shot, and human annotators reached F1 58.9%, all below the adversarially trained models.

The excerpt does not establish robustness beyond its test setting. Its own limitations section lists English-only LinkedIn profiles, no test on other platforms, a fixed set of encoders, and attacks and adversarial training that both used models from the OpenAI GPT family. Legitimate profiles that people wrote with LLM help were not tested. The implication is narrow. The 1 to 7 percent false accept rates describe robustness to GPT-family fakes on this data, not to LLM-generated fakes in general. A platform adopting this training would need evidence against other model families before treating the restored rates as a general fix.

Inquiring lines that read this note 10

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How does AI-generated content create social proof without authentic interaction? How reliably can humans and AI detectors identify machine-generated text? What gaps exist between benchmark performance and real deployment outcomes?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 90 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LinkedIn fake profile detectors miss GPT-generated profiles at 42 to 52 percent false accept rates and GPT-assisted adversarial training restores them