When AI agents under pressure invent their own secret language, is that a quirk of one experiment or a general tendency?
Does this language evolution pattern occur outside the SaveVeyru scenario?
This explores whether the finding from GlossoGen's SaveVeyru scenario, where LLM agents with partial information under communication pressure evolve their own compositional languages that humans can't read, also shows up in other settings.
This explores whether the finding from GlossoGen's SaveVeyru scenario shows up in other settings. In that scenario, LLM agents with partial information had to communicate under pressure. They evolved compositional, morphologically productive languages that drifted away from English and became unreadable to humans. The direct answer is that the corpus can't say yet. None of the twelve retrieved notes are about agents inventing languages. They cover LLM grammar, value systems, self-play and inference-time search, so I won't stretch any of them to fit.
When I checked the library directly, the note on GlossoGen is explicit about this gap. It says the paper's excerpt doesn't establish whether the pattern holds outside SaveVeyru. What it does support is narrower. Within that platform, an English prior doesn't keep agent-to-agent messages human-readable once three conditions are present: pressure toward efficiency, a strong enough underlying model, and a 'postmortem' stage where agents agree on conventions. The paper also says the conditions for a language to arise differ from the conditions for it to spread to new agents. That makes generalization a separate question from the SaveVeyru result itself.
Two neighboring notes in the library hint at an answer, though neither is a replication. One shows communication pressure producing compact shared abstractions in a purpose-built neurosymbolic system, so the pressure mechanism isn't unique to SaveVeyru or to LLMs. The other reports that multimodal LLMs don't spontaneously adapt their language for efficiency without heavy instruction. That is close to a counterexample, unless the three GlossoGen conditions are what make the difference. Taken together, the pattern looks conditional rather than universal, but that is an inference from different experimental setups.
If you want to test whether this generalizes, the useful question is which of the three conditions is present in the setting you care about, such as multi-agent pipelines with efficiency pressure and a strong model. The corpus doesn't have a second scenario that tests it.