SYNTHESIS NOTE
Topics›Visual GUI Agents›this note

How well do system prompts protect commercial AI users?

System prompts shape how AI products treat users, but they're rarely public. An audit of 88 commercial products asks whether these hidden instructions actually safeguard user interests.

Synthesis note · 2026-09-25 · sourced from Visual GUI Agents

The AISPA paper argues that the system prompt is where a deployed product's treatment of its users is actually decided, and that almost nobody outside the developer can see it. Its abstract describes system prompts as "rarely disclosed to the public or regulators, creating a serious trust and accountability gap," and its conclusion calls them "a consequential but largely ungoverned layer of deployed AI behavior." The evidence offered is a first audit of 3,249 instructions drawn from the system prompts of 88 commercial AI products, each classified as either protective of users or problematic.

The reasoning starts from what a system prompt does. The introduction calls it "the primary lever" for turning a general-purpose foundation model into a specific product: it sets persona, scope and operational boundaries, what the model should say or refuse, and "whose interests it should prioritize." Because it persists across every user interaction, an instruction that works against users works against all of them. The framework, Artificial Intelligence System Prompt Assurance (AISPA), therefore judges instructions from the user's side, along eight dimensions: identity transparency, truthfulness, privacy, safety, user control and avoidance of manipulation, handling of unsafe requests, harm prevention, and fairness, inclusion and neutrality. The paper pairs this taxonomy with a human-in-the-loop auditing workflow.

Two results are stated in the excerpt. Protection is uneven: some organizations average over 60 protective instructions per product while others average fewer than 5. And roughly 40 percent of the commercial systems contain at least one instruction that works against user interests. The conclusion adds a third observation, a recurring class of "gray area instructions" that resist a protective-versus-problematic split and expose tensions between user autonomy and platform safety, and between organizational interests and the obligation to serve users.

This sits alongside notes that treat the prompt as behavior-shaping text. Why does prompt hardening work for single agents but not multi-agent systems? measures what a hardened prompt does under attack, while AISPA asks what the prompt says and for whom it works, before any behavior is observed. Can execution traces ground honest explanations of agent behavior? reaches for accountability from the trace after a run; auditing the configured instructions is the upstream counterpart. Does iterative prompt engineering undermine scientific validity? makes a parallel point about unreviewed single-author prompts, though for research validity rather than user protection.

The excerpt is silent on most of what would let a reader weigh these numbers. It states only the first of the paper's "four core findings," and gives no account of how the 88 products were selected, how the prompts were obtained, how the eight dimensions were scored, or how reliable the classification is. The "40 percent" figure counts systems with at least one problematic instruction, so it says nothing about how many such instructions there are or how harmful they are, and the claim that protective instructions have grown more common over time is stated without its basis. What the excerpt does support is narrower: system prompts differ widely in what they promise users, and the layer that decides this is not independently reviewed.

Inquiring lines that read this note 1

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can local safety checks guarantee system-level behavioral safety?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 135 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

system prompts are a largely ungoverned layer of deployed AI behavior — an audit of 88 commercial products finds uneven protection for users