INQUIRING LINE

If you train a fake-detector on GPT's fakes, does it catch fakes from Claude, Llama, or Mistral too?

Does adversarial training on GPT fakes work against other LLM families?

This explores whether a fake-content detector that learns to catch text from one AI model family (here, GPT) would also catch fakes written by other families like Llama, Claude or Mistral. The corpus shows the training works for GPT, but it has no direct test of whether that carries over to other families.


This explores whether a fake-content detector that learns to catch text from one AI model family (here, GPT) would also catch fakes written by other families like Llama, Claude or Mistral. The corpus shows the training works for GPT, but it has no direct test of whether that carries over to other families. The clearest evidence comes from LinkedIn fake-profile detection Can fake profile detectors catch GPT-generated LinkedIn profiles?. Detectors trained on real profiles and hand-written fakes let 42–52% of GPT-written fakes through. Retraining them on GPT-generated examples cut that to 1–7%, and genuine profiles weren't flagged any more often. That shows machine-written fakes are a different kind of fake from human-written ones, and a detector has to see them to catch them. It doesn't show what happens when the fakes come from a model it never saw.

Other notes suggest why you shouldn't assume the protection carries over. Language models leave model-specific fingerprints. Each model tends to invent its own recurring sets of fictional names, and those patterns survive paraphrasing and copying well enough to trace more than 1,600 ghost-written papers back to their source Do language models leak their training through fictional names?. This cuts both ways. A detector trained on GPT fakes may be learning GPT's particular habits rather than anything true of AI writing in general. If so, it could do very well on GPT and miss a Llama or Mistral fake that has different habits. The detector would be learning a style, not the general fact that a machine wrote the text.

The opposite case is also in the corpus: some weaknesses really do carry over between model families. Persuasion-based jailbreaks worked more than 92% of the time on GPT-3.5, GPT-4 and Llama-2 alike Can social science persuasion techniques jailbreak frontier AI models?. AI judges from different families fall for the same tricks, such as fake citations and polished formatting, because those tricks work at the level of surface signals rather than meaning Can LLM judges be fooled by fake credentials and formatting?. The useful takeaway is that detection carries over only when it latches onto something every model family shares. It breaks when it latches onto one family's habits.

Training one model against another adversarially is a possible way forward. In one study, a critic model learned to tell expert answers from a model's own answers, and it worked across very different tasks Can adversarial critics replace task-specific verifiers for reasoning?. In another, two models debating in front of a weaker judge stopped training from gaming that judge's blind spots Can debate training prevent reward hacking by weaker judges?. Neither is about detecting fakes, but both point to the same idea: a detector that keeps being challenged by new generators is less likely to overfit to one of them. To answer the question properly, someone would need to train on one family's fakes and test on another's, and the corpus doesn't contain that experiment yet.


Sources 6 notes

Can fake profile detectors catch GPT-generated LinkedIn profiles?

Detectors trained on genuine and manual fakes miss GPT-generated profiles at 42–52% false accept rates, but adversarial training on GPT-generated data restores detection to 1–7% false accepts without raising false rejects.

Do language models leak their training through fictional names?

LLMs generate non-random, model-specific name combinations that act as behavioral fingerprints and survive across paraphrase and copy. Over 1,600 ghost-authored papers with valid DOIs show this leakage has already contaminated scholarly infrastructure.

Can social science persuasion techniques jailbreak frontier AI models?

A 40-technique taxonomy of psychology-based persuasion strategies (PAP) achieved over 92% attack success on GPT-3.5, GPT-4, and Llama-2 in 10 trials. Current defenses miss semantic content attacks because they screen for unusual patterns, not fluent persuasion.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Can adversarial critics replace task-specific verifiers for reasoning?

RARO uses an adversarial game where a critic discriminates expert from policy answers, eliminating the need for domain-specific verifiers while matching the scaling properties of verifier-based RL. The approach works across Countdown, DeepMath, and Poetry Writing tasks.

Show all 6 sources
Can debate training prevent reward hacking by weaker judges?

On math tasks, debate between a generator and critic adjudicated by a frozen weaker judge maintained judge performance throughout training and achieved 45% higher peak validation accuracy than single-player RLAIF, which quickly exploited the judge's errors and collapsed in accuracy.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.