INQUIRING LINE

When five teams each define "active user" differently, whose definition wins — and who's actually in charge of deciding?

How do organizations govern metric definitions across multiple teams and systems?

This explores how organizations keep a metric meaning the same thing when many teams and systems define, compute and use it. The collection has little on that business-data question directly, so this answer draws on nearby research about shared rules, agreement and measurement across multi-agent AI systems.


This explores how organizations keep a metric meaning the same thing when many teams and systems each define, compute and use it. To be direct: the collection has nothing on the usual tools for this, such as semantic layers, data catalogs, metric stores or data stewardship. What it does have is AI agent research that runs into the same governance problem in a different setting, and those papers show where metric governance tends to break.

The first lesson is about ownership. Work on agents that hand tasks across organizational boundaries finds that nobody is named as the owner of the rules those tasks must follow Who enforces invariants when agents cross organizational boundaries?. Rules can come from the operator, the organization, a regulator or a standards body. Their policies may conflict, and not every party can see all of them. Metric definitions fail the same way. When 'active user' passes through three teams' pipelines, the real question is less *what* the definition is than *whose* definition wins when they disagree, and whether everyone can even see the competing versions.

The second lesson is that agreement and correctness are different guarantees. A study of validator consensus shows that a protocol can force everyone to agree, but it can only make it statistically likely that what they agree on is actually right Can validator consensus guarantee both agreement and semantic correctness?. A governance process that gets every team using the same metric name has solved the first problem, not the second. Other research shows that agents accept information from their neighbors without checking it, so errors spread quietly through the network Why do multi-agent systems fail to coordinate at scale?. A flawed upstream definition travels the same way.

The third lesson is about rollout. Coordination standards spread by wrapping and connecting existing protocols, not by forcing everyone to replace them Should coordination protocols wrap existing systems or replace them?. Read across to metrics, that suggests a governance layer that maps each team's existing definitions to a shared one will be adopted faster than a mandated rewrite. There is a counterweight, though. Production teams found that flexible protocol-based integration caused unpredictable failures, and that explicit, fixed function calls restored reliability Why do protocol-based tool integrations fail in production workflows?. Bridging helps adoption, but the definitions themselves need to be strict and unambiguous.

A final, less comfortable point: a single headline number can hide what matters. Agent evaluation research shows that identical success rates can hide very large differences in efficiency and reliability How should we measure agent system performance beyond task success?. Part of governing a metric is deciding what it is allowed to hide. If you want literature on enterprise metric governance itself, this collection is not the place yet.


Sources 6 notes

Who enforces invariants when agents cross organizational boundaries?

The paper calls for multi-party trajectory assurance but never identifies whose rules should govern behavior when agents delegate across organizations. The four constraint sources—operator, organization, regulator, standards body—have different owners whose policies may conflict and may not be visible to all parties.

Can validator consensus guarantee both agreement and semantic correctness?

Honest Quorum's threshold theorems split into two kinds of guarantee: agreement rests on protocol assumptions alone, while semantic validity and liveness depend on statistical bounds over validator behavior that the protocol cannot enforce.

Why do multi-agent systems fail to coordinate at scale?

AgentsNet benchmark shows agents fail to coordinate strategies either by agreeing too late or adopting strategies without informing neighbors. Agents accept neighbor information without verification, enabling error propagation while remaining capable of detecting direct conflicts.

Should coordination protocols wrap existing systems or replace them?

Research shows that agent coordination standards achieve adoption by composing existing protocols like MCP and DIDComm under a shared substrate, rather than competing to replace them. Bridging lets value accrue incrementally without forcing ecosystem-wide rewrites.

Why do protocol-based tool integrations fail in production workflows?

MCP integration caused non-deterministic failures through ambiguous tool selection and parameter inference. Replacing it with explicit direct function calls and single-tool-per-agent design restored determinism. A 306-practitioner survey confirms 85% of production teams build custom agents, forgoing frameworks.

Show all 6 sources
How should we measure agent system performance beyond task success?

Single task-success metrics obscure how agents achieve results across memory, context, and verification layers. Research shows identical success rates can mask enormous differences in efficiency, reliability, and deployment readiness—requiring harness-level benchmarks that measure trajectory, memory hygiene, and verification costs.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.