Weak Links in LinkedIn: Enhancing Fake Profile Detection in the Age of LLMs
Abstract. Large Language Models (LLMs) have made it easier to create realistic fake profiles on platforms like LinkedIn. This poses a significant risk for text-based fake profile detectors. In this study, we evaluate the robustness of existing detectors against LLM-generated profiles. While highly effective in detecting manually created fake profiles (False Accept Rate: 6 −7%), the existing detectors fail to identify GPT-generated profiles (False Accept Rate: 42−52%). We propose GPT-assisted adversarial training as a countermeasure, restoring the False Accept Rate to between 1 −7% without impacting the False Reject Rates (0.5 −2%). Ablation studies revealed that detectors trained on combined numerical and textual embeddings exhibit the highest robustness, followed by those using numerical-only embeddings, and lastly those using textual-only embeddings. Complementary analysis on the ability of prompt-based GPT- 4Turbo and human evaluators affirms the need for robust automated detectors such as the one proposed in this study.
Introduction. Online professional networks, such as LinkedIn, play a crucial role in professional interactions, hosting over 1.15 billion active users and generating significant economic activity [1]. However, such platforms face growing threats from fake profiles used for phishing, misinformation, and recruitment fraud [2,3]. Recent advances in Large Language Models (LLMs), particularly GPT-3.5 and GPT-4, have simplified the creation of highly realistic fake profiles, posing a significant threat to the existing detectors [4, 5]. Between 2021 and 2022, the number of fake profiles on LinkedIn nearly doubled [5]. Prompt-based evaluations of humans and GPT-4 achieved modest detection accuracy (F1 Human: 59%, F1 GPT-zero shot: 71%, F1 GPT-few shot: 86%, see Section 4.3). Existing detection approaches, such as Section and Subsection Tag Embeddings (SSTE) proposed in [4], perform well (F1∼96%) against manually created fake profiles but fail sharply (F1∼68%) against LLM-generated profiles. To address these challenges systematically, we pose and address q1 How vulnerable are current detection methods to profiles generated by advanced LLMs? q2 Can adversarial training with LLM-generated profiles enhance detection robustness? q3 How does the effectiveness of our proposed detection methods compare to human evaluators and GPT-4? Our primary contributions are as follows. 1 We implement a series of robust fake profile detection systems using textual, numerical, and fused features, incorporating Section Tag Embeddings (STE), Section and Sub-Section Tag Embeddings (SSTE), along with PCA-based dimensionality reduction. The best setup outperformed prior methods on genuine and manual fake profiles [6]. 2 We augment the existing dataset with 600 fake profiles that we generated using GPT- 4-Turbo with carefully crafted prompts. These synthetic profiles closely mimic legitimate users, as verified by similarity metrics, and used for creating attack vectors. 3 We demonstrate that existing detectors fail against LLM-generated profiles (FAR: 42−52%) and proposed GPT-assisted adversarial training, which restored FAR to 1−7% without compromising the legitimate user classification.
4 We also benchmark detection capabilities of human annotators and GPT-4, confirming the need for ML-based automated detectors. 5 Finally, we conducted ablation studies revealing text embeddings alone are fragile under LLM attack, whereas numerical profile features remain sturdier; their fusion yields the most robust detector. The remainder of this paper is structured as follows: Section 2 reviews related literature, Section 3 describes materials and methods, Section 4 reports and discusses results, Section 5 limitations and future research directions, and Section 6 concludes.
Related work. Research on fake profile detection spans numerical, graph-based, behavioral, and textual methods. Early efforts used correlation-based analysis of profile attributes. For instance, Adikari et al. [3] achieved 87.34% accuracy on LinkedIn profiles; however, their approach relied on historical data and assumed attribute consistency, which are limitations when handling cold-start accounts. Graphbased models, such as SybilBelief [7], SybilEdge [8], and SybilFlyover [9], leverage network topology and user connectivity, often achieving AUCs above 0.9. However, they require relational metadata (e.g., connections, followers), limiting their applicability for newly created or minimally active profiles. Early stylometric techniques relied on N-grams and writing patterns [10]. The LLM-assited fake profile detection problem is similar to LLM-assisted cheating detection [11,12]. Recent keystroke dynamics-based approaches [13, 14] achieve a detection accuracy close to 95%. Using activity-based features such as post frequency and follower-following ratios, Alnagi et al. [15] employed XGBoost with SHAP-based interpretability, achieving 94% precision on Instagram and 91% on Twitter. Ayoobi et al. [4] proposed Section and Subsection Tag Embeddings (SSTE), reaching 95% accuracy on genuine and manually crafted fake LinkedIn profiles. However, their performance drops sharply (to 71–76% accuracy) against GPTgenerated profiles, revealing a growing vulnerability. Our work differs from Ayoobi et al. [4] and focuses on (1) investigating the robustness of baseline detectors (trained on genuine and manually created fake profiles) against profiles generated by GPT3.5 and GPT4Turbo, (2) assessing the power of GPT3.5, GPT4Turbo, and GPT3.5+GPT4Turbo-assisted adversarial training, and (3) evaluating the performance of Human and GPT-based evaluators systematically.
Method. We trained models on LLP vs FLP as a baseline, and introduced three attack and three adversarial training scenarios using GPT3.5Ps, GPT4Ps, or both. Each model was evaluated on all four profile types (LLPs, FLPs, GPT3.5Ps, and GPT4Ps). Classifiers were trained using STE embeddings from RoBERTa, DeBERTa, ModernBERT, and Flair with XGBoost and CatBoost. Models were evaluated using F1 score, false accept rate (FAR; fake →legitimate), and false reject rate (FRR; legitimate →fake). The effectiveness of the attack and countermeasure was measured by changes in FAR under adversarial conditions and after retraining. Calibration was assessed using reliability curves [30], which compare predicted probabilities to empirical frequencies; the diagonal indicating perfect calibration. Brier score [31] was also computed as a scalar measure of calibration, where lower values indicate more reliable confidence estimates, critical for minimizing overconfident mis-classification of LLM-generated profiles.
3.5 Human and GPT-4 evaluation We benchmarked human and GPT-4 detectors on the same inputs: Name, Location, Education, Experience, Skills, Connections, Followers, Summary, and derived statistics. GPT-4 was tested on 360 profiles (180 real, 180 fake) using the
Discussion. Vulnerability to GPT-generated profiles: On GPT3.5+4P attacks, F1 scores dropped to 67.88%–73.82%, and FARs rose to 52.1% (DeBERTa+CatBoost), indicating over half of sophisticated fake profiles were misclassified as legitimate. This degradation aligns with high textual similarity between GPT-generated and real profiles (mean: 88.9%, range: 64.2%–99.4%).
Adversarial training: GPT3.5-assisted training improved F1 to 97.83%–98.15% and cut FARs on GPT3.5Ps to as low as 1.3%, but remained vulnerable to GPT4Ps (FARs: 16.9%–19.3%). GPT4-assisted training reversed this—F1 up to 97.84% and FARs on GPT4Ps down to ∼2%, but showed limited generalization to GPT3.5Ps. In contrast, training on the combined GPT3.5+4P dataset yielded strong generalization across all attacks, with FARs between 1.34% and 2.6% and F1 scores consistently above 97.5%. Flair+XGBoost achieved the best overall performance (F1 = 98.2%, FAR = 1.34% on combined attacks). Across all adversarial training settings, FRRs remained stable (1.48%–2.41%), confirming no significant compromise on correctly classifying legitimate profiles.
Figure 3 compares the performance of human evaluators and GPT-4 on the task of LinkedIn fake profile detection. Human evaluators showed limited effectiveness, particularly on GPT-generated profiles, with a three-class accuracy of only 31.4% on this category. Aggregated into a binary classification task, their F1 score was 58.9%, with a false accept rate (FAR) of 38.7% and a false reject rate (FRR) of 46.6%, indicating considerable confusion between real and fake profiles. GPT-4 performed better overall. In the zero-shot setting, it reached an F1score of 71.3% but misclassified 43.9% of fake profiles. With few-shot prompting, its performance improved substantially, achieving perfect accuracy on legitimate profiles and reducing the FAR on fakes to 25.0%, raising the F1 score to 85.7%. However, both approaches fall short of our adversarially trained models, which consistently achieve F1 scores above 97.5% and FARs below 2% across all LLMgenerated profile scenarios.
Text vs. numerical features (post-adversarial training) Feature types also responded differently to adversarial training. Using only STE embeddings, GPT4Passisted training reduced FARs to 2.59%–8.33% (GPT3.5Ps) and 3.12%–8.33% (GPT4Ps), indicating improved generalization across LLM variants. In contrast, models using numerical features demonstrated asymmetric gains. For the full 167-dimensional feature set, GPT4P-assisted training substantially reduced FARs for GPT4Ps but had a limited impact on GPT3.5Ps. For example, Flair+XGBoost: FAR dropped from 38.52% to 36.11%. A similar trend was observed when using only numerical features: GPT3.5P-assisted training improved F1 scores from 78.67%–79.21% to 82.55%–84.08% on GPT3.5Ps, whereas GPT4P-assisted training led to a larger jump for GPT4Ps (up to 96.75%). These findings suggest that textual features tend to generalize more effectively across model variants and adversarial scenarios, whereas numerical features seem to encode generation-specific artifacts. Moreover, the compact 17- dimensional numerical representation offers a lightweight alternative for detection in resource-constrained settings.
Conclusion. Existing LinkedIn fake profile detectors perform well on manually created profiles (F1 > 95%) but fail on GPT3.5 and GPT4-generated ones, with F1 dropping to 67.88% and false accept rates (FAR) exceeding 52%. Human annotators (F1 = 58.9%) and general-purpose LLMs (F1 = 85.7%) also underperform in this setting. Targeted adversarial training using GPT-generated profiles restored F1 to 98.2% and reduced FAR to 1.34%, with minimal impact on legitimate profile rejection (FRR < 2.5%). Flair embeddings with XGBoost gave the most consistent results. Ablation experiments revealed that textual features degrade sharply under attack, while numerical features remain more robust. Their combination yields better generalization across model variants and input conditions. These findings support the need for task-specific retraining to maintain robustness against high-quality synthetic profiles generated with the help of LLMs.
Limitations. While our approach achieves low FARs (1.34%–2.28%) through STE-based features, PCA, and adversarial training under both adversarial and non-adversarial environments, several limitations remain. First, the evaluation is restricted to English-language LinkedIn profiles and should be extended to languages other than English. The method’s applicability to other platforms or multilingual contexts has not been tested. Second, embedding extraction relies on a fixed set of LLM-based encoders (e.g., DeBERTa, RoBERTa), which may affect stability as model architectures evolve. Third, both the creation of fake profiles for generating attack vectors and adversarial training used models from the same LLM family (OpenAI GPT). The system needs to be evaluated against a wide variety of advanced LLMs. Fourth, the current experiments should be expanded to include legitimate profiles created by legitimate people who utilized LLMs for creating and polishing their profiles.
Lines of inquiry this paper opens 17
Research framings built by reading the notes related to this paper — the questions it feeds into.
How does AI-generated content create social proof without authentic interaction?- How often does LinkedIn wrongly flag legitimate posts as AI-generated?
- How similar are GPT-generated fake profiles to real human profiles?
- How does LinkedIn's verification system affect what content appears in feeds?
- Can platforms trust these detection rates on profiles created with human-LLM collaboration?
- What triggers LinkedIn's detection of inauthentic content from heavy AI use?
- How does LinkedIn's platform response address detected AI-generated content?
- How much of LinkedIn's feed is genuinely AI-generated versus human-written content?
- How accurate is the detector labeling these posts?
- How does LinkedIn's approach differ from other AI content moderation systems?
- How does LinkedIn's comment-versus-post AI split compare to Reddit's?