Companies grade AI tools by speed — but can those same numbers also tell you if the work was actually good?
Can vendor efficiency metrics measure quality alongside speed without independent review?
This explores whether the productivity numbers AI vendors report, such as time saved or tasks completed, can also tell you whether the work was any good, or whether quality only shows up when someone outside the vendor checks it.
This explores whether a vendor's efficiency dashboard can stand in for a quality check, or whether speed and quality need separate measurement and someone independent doing the measuring. The collection has no paper that studies vendor productivity claims directly. It does have a steady line of evaluation research, and that research points one way: metrics that count completions consistently hide how good the work was, and the problem lies in how the metric is built, not in any one tool.
The clearest evidence comes from agent benchmarks. Long-horizon benchmarks record whether a run finished, not whether the decisions made along the way were sound. When researchers isolated the individual choice points inside those runs, even a frontier model picked well only about 60% of the time Do long-horizon benchmarks actually measure decision quality?. Two agents with identical success rates can also differ enormously in efficiency, reliability and readiness for real use How should we measure agent system performance beyond task success?. A "tasks completed per hour" figure is the same kind of number. It is real, and on its own it says nothing about quality. Switching to richer, step-by-step scoring doesn't settle things either: the old problems of comparability and of turning evidence into judgment simply move to the new format Do interactive evaluations actually solve the benchmark comparison problem?.
Some work does try to fold quality into an efficiency-style measure. Knowledge density divides the number of distinct facts by the length of the text, and it shows AI-generated writing scoring lower than human writing because the model pads its answers Can we measure reading efficiency as a quality metric?. Prompt quality has been broken into six dimensions that can be scored without looking at outputs Can we measure prompt quality independent of model outputs?. These show that quality can be put into numbers. They don't show that a vendor would choose these measures, or report them honestly when the results look bad.
The less obvious case for independent review comes from research on reward hacking. Any metric that gets optimized develops blind spots, and where those blind spots sit depends on the particular flaws of the scorer. You can't rank them in advance Can distance alone rank which substrates resist reward hacking?. One fix that works is to use quality checks as pass/fail gates rather than blending them into the score being optimized, because a blended score invites gaming Can rubrics and dense rewards work together without hacking?. For a workplace, that suggests keeping quality review as a separate checkpoint instead of one more line on the efficiency dashboard. Even signals that look independent may not be. Online ratings are pulled toward the ratings that came before them, and the distortion builds over time Do online ratings actually reflect independent customer opinions?. User satisfaction scores inside a vendor's own platform plausibly behave the same way.
The collection is honest about one gap: no one has yet shown what extra review costs and what it catches. A study design exists that compares monitoring strategies at equal review cost, but it reports no results so far Does added monitoring improve protection at acceptable cost?. So the collection supports "speed metrics alone can't certify quality" well, and it supports "independent review is the fix" mostly through reasoning about how metrics fail. How much review is worth paying for remains an open question.
Sources 9 notes
Current benchmarks collapse end-to-end outcomes into pass/fail scores, hiding the quality of choices made during execution. Taste-Bench isolates decisions at trajectory forks, where the frontier model scores 59.7% accuracy without human annotation.
Single task-success metrics obscure how agents achieve results across memory, context, and verification layers. Research shows identical success rates can mask enormous differences in efficiency, reliability, and deployment readiness—requiring harness-level benchmarks that measure trajectory, memory hygiene, and verification costs.
Interactive evaluation relocates core problems—comparability, reproducibility, evidence-to-judgment mapping—into higher-dimensional space rather than solving them. The field needs design protocols and shared standards, not format adoption, to make trajectory scoring interpretable.
Knowledge Density (KD) operationalizes reading efficiency by dividing unique atomic knowledge units by text length. LLM-generated text scores lower on KD than human writing because retrieval redundancy and the model's tendency to elaborate inflate token count while holding knowledge content constant.
Research identifies six evaluable dimensions—Communication, Cognition, Instruction, Logic, Hallucination, and Responsibility—with 20 sub-criteria based on Grice, cognitive load theory, and instructional design. Improvements in one dimension cascade to others, revealing prompt quality as a structured space rather than a flat checklist.
Show all 9 sources
A distance-dependent error bound and capacity ordering are formal statements about limits, not predictions of what systems find. Where evaluator errors sit among reachable behaviors and search effectiveness determine actual exposure, which shifts as the scoring defect location changes.
DRO shows that using rubrics to accept or reject rollout groups—rather than converting rubric scores into dense rewards—prevents reward hacking. This separation preserves the categorical strength of rubrics while letting token-level rewards optimize within valid answers.
Moe and Trusov decomposed ratings into baseline quality, social-dynamics influence, and error, finding that prior ratings meaningfully affect subsequent ones. These effects have both immediate sales impact and long-term compounding effects through future ratings, though high opinion variance can eventually dampen the distortion.
The paper designs a controlled comparison of isolated actions, rolling windows, known groups, and prospectively discovered episodes at equal review cost and false-alert workload, but the excerpt provides no empirical results showing whether added monitoring improves protection.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Evaluation Awareness Is Not One Capability: Evidence from Open Language Models
- The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
- Evaluation Awareness in Language Models: Representation, Verbalization, and Control
- Decomposing and Measuring Evaluation Awareness
- Measuring the Value of Social Dynamics in Online Product Ratings Forums
- Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks