INQUIRING LINE

Could an AI's answers to a personality questionnaire be enough to tell you which chatbot wrote them?

Can model identity be recovered from psychometric response patterns alone?

This explores whether the way a model answers standardized psychological questionnaires works like a fingerprint, so that its answers alone would tell you which model produced them.


This explores whether a model's answers to standardized psychological questionnaires work like a fingerprint that tells you which model you're talking to. The corpus says the fingerprint is real, but it hasn't been tested as a true "guess the model" game. It also lives mostly in the leftovers, not in the headline scores.

The strongest evidence is a study that gave nine LLMs seven psychological instruments, in both Chinese and English. Each model showed a stable, model-specific response profile that reproduced across repeated administrations Do LLMs show reproducible psychological profiles when given standardized tests?. Identification needs exactly that: answers that are consistent within one model and different between models. But all nine models also converged on the same prosocial, stability-oriented pattern, which is presumably the imprint of shared alignment training. That shared core carries almost no identifying information, so whatever separates the models sits in the differences around it. The note establishes that the profiles are reproducibly distinct. It doesn't report how accurately a classifier could pick the model out of a lineup.

The most surprising finding is that what a model declines to answer is part of its signature. Structured non-response patterns distinguished the models too, meaning each one draws its own boundary around which questions count as applicable to it Do LLMs show reproducible psychological profiles when given standardized tests?. Scoring only the numbers a model gives would throw away some of the most identifying evidence.

The reason a questionnaire can carry identity at all is that the answers needn't be true. LLM self-reports mostly echo the distribution of human text they were trained on, not a look inside the model Can language models actually introspect about their own states?. So a psychometric profile fingerprints training and alignment, not an inner psychology. Fingerprinting only needs consistency. A related result points the same way: models pass behavioral traits to each other through data with no semantic connection to the trait. That effect is model-specific and fails across different architectures, so it embeds statistical signatures below the level of meaning Can language models transmit hidden behavioral traits through unrelated data?. It's a different channel, but the lesson is the same: identity leaks through places that look like content but aren't.

The open question is how movable the fingerprint is, and the corpus doesn't test it directly. Persona prompts make models follow trait instructions at the output level while leaving underlying bias unchanged Can persona prompts actually reduce bias in language models?. That suggests a system prompt could shift questionnaire scores while something deeper stays put. Persona traits also correspond to directions in activation space that shift during finetuning Can we track and steer personality shifts during model finetuning?. So a finetuned descendant may drift away from its base model's profile. Both are my inferences, not tested results. The defensible answer today is that psychometric responses are distinctive enough to work as a signature, but the corpus has no robustness data on prompting or finetuning, and no measured identification accuracy.


Sources 0 notes