Are the number-based tells used to spot AI writing really fingerprints of one model, like GPT-4, or of AI text in general?
Do numerical features encode artifacts specific to GPT-3.5 or GPT-4?
This explores whether the measurable numeric signals people extract from AI-generated text (such as statistical or stylistic features used to detect it) pick up quirks of one particular model like GPT-3.5 or GPT-4, rather than traits of AI writing in general.
This explores whether the numeric signals used to describe or detect AI output capture quirks of one particular model, such as GPT-3.5 or GPT-4, rather than something true of AI text in general. The short answer is that this collection doesn't directly address it. None of the retrieved notes test feature-based detectors across GPT versions or check whether a detector trained on one model's text carries over to another. Instead of padding, here is what the corpus does offer: indirect evidence that models leave fingerprints tied to their capability tier and their training history, which is the premise your question depends on.
The clearest case is that models fail in recognizably different ways depending on how capable they are. When LLMs repeatedly edit documents, weaker models visibly delete content, while frontier models quietly corrupt it and leave the surface looking intact Does model capability change how documents degrade?. That is a model-tier artifact: a pattern you could measure that shifts as you move from a GPT-3.5-class model to a GPT-4-class one. It also suggests a trap. Any feature tuned to catch the obvious errors of older models may miss the subtler errors of newer ones.
A second source of model-specific artifacts is training. RL post-training tends to amplify one output format the model saw during pretraining and suppress the others. Which format wins depends on model scale, not on which format works best, and this is largely hidden when you start from a proprietary model Does RL training collapse format diversity in pretrained models?. So two models from the same vendor could settle on different 'house styles' for reasons outsiders can't see, and those styles are exactly what numeric features would end up encoding. Model-specific blind spots also persist at the level of content. GPT-4o's image generation reliably distorts non-Latin scripts in ways that trace back to its training data Why do unified image generators fail on non-Latin scripts?. LLMs also produce plausible-looking but wrong numbers instead of actually running iterative calculations, a failure that holds across model scales Do large language models actually perform iterative optimization?.
The surprise for a curious reader is that fingerprints come in two kinds, and they behave differently. Some, like the pattern-matched numbers, persist across model generations and might support a general detector. Others, like how documents degrade or which formats RL settles on, shift with each model's scale and training, so a detector built on GPT-3.5 output may not survive the move to GPT-4. To answer the original question rigorously, the collection would need papers on cross-generator detection or stylometry, which it currently lacks.
Sources 4 notes
DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.
Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.
GPT-4o demonstrates impressive unified multimodal generation but systematically produces distorted or Latin-approximated characters for non-Latin scripts and underrepresented cultures. This pixel-space failure directly reflects training data over-representation of Western languages and cultures.
Research shows LLMs cannot perform iterative procedures in latent space. They recognize optimization problems as template-similar and emit plausible-looking but incorrect values, a failure mode that persists across model scale and training approaches.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- An Empirical Study of GPT-4o Image Generation Capabilities
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- The Art of Scaling Reinforcement Learning Compute for LLMs
- Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
- LLMs Corrupt Your Documents When You Delegate
- Sharpening Tax in Post-Training
- Can Large Language Models Reason and Optimize Under Constraints?