Measuring What Really Matters
The AI industry is over-indexing on benchmarks and under-measuring what really matters.
The AI industry is getting better at showing agents doing things. It is still much weaker at proving that those things were valid, useful, authorized, representative, or safe.
That is the problem underneath three recent research threads: Mnemosyne, automated benchmark auditing, and METR’s Frontier Risk Report. The issue is not simply that AI measurement needs to “catch up.” It is that many of the metrics are skewed. They measure convenient slices of performance and then treat those slices as a picture of reality.
What Matters Today
Mnemosyne: A generated agent action can look valid while still being stale, infeasible, conflicting, or destructive of the evidence that triggered it. The paper proposes Agentic Transaction Processing (ATP), a transaction model for generated workflows under which a generated action must earn transaction authority before it becomes committed truth, logs, admission rules, repair, compensation, and evidence preservation before actions are treated as complete. As the authors put it, what is new is not the failure but the proposer: capable AI now drafts and repairs workflows at a speed and scale no human author matches. (arxiv.org)
Automated Benchmark Auditing: The paper argues that flawed benchmark tasks distort capability assessments. The auditing agent identified ambiguous task design, execution environment conflicts, and incorrect ground truths in over 25.7% of evaluated tasks, and filtering out the problematic tasks shifted model rankings and increased average performance by 9.9% on SWE-bench Verified and 9.6% on Terminal-Bench 2. This is not an isolated finding. OpenAI itself stopped using SWE-bench Verified for frontier launches after auditing a subset of frequently failed problems and finding that at least 59.4% of them had flawed test cases that rejected functionally correct submissions. The industry’s most-cited coding benchmark was, in meaningful part, grading against broken ground truth. (arxiv.org; openai.com)
METR Frontier Risk Report: METR conducted a pilot exercise to assess misalignment risks from AI agents used inside frontier AI developers, with participation from Anthropic, Google, Meta, and OpenAI. The assessment concluded that agents in February–March 2026 plausibly had the means, motive, and opportunity to launch a minimal “rogue deployment,” but lacked the means to make rogue deployments robust to serious efforts to shut them down. METR expects the plausible robustness of rogue deployments to increase substantially in the coming months. (metr.org)
The Signal
The industry is mistaking measurable activity for trustworthy work.
An agent completed a task. But was the task representative?
A benchmark score improved. But was the benchmark clean?
A workflow ran automatically. But were the actions authorized, reversible, and supported by preserved evidence?
A model performed well in a test environment. But did the test environment reflect the messy, permissioned, adversarial, and operational reality where the system will actually be used?
That is the gap. Not “measurement needs to catch up.” More pointedly: the measurements being used are often too narrow, too convenient, and too disconnected from real-world verification.
John Doerr Named This Decades Ago
None of this is a new failure mode. It is the oldest one in management, and John Doerr wrote the book on it, Measure What Matters, drawing on the Objectives & Key Results “OKR” system Andy Grove built at Intel and Doerr later carried to Google.
Two of Grove’s disciplines map directly onto the AI industry’s current predicament.
First, measure output, not activity. Grove’s foundational insight was that busy organizations routinely confuse effort with results, lots of motion, little output. A benchmark score is an activity metric dressed up as an output metric. It tells you the model did something in a controlled slice of reality. It does not tell you whether real work product was produced, whether it held up, or whether anyone should have relied on it.
Second, key results must be verifiable. In Doerr’s framework, a key result is worthless unless you can look at it later and prove, without argument, that it happened. As Marissa Mayer’s rule at Google put it: “It’s not a key result unless it has a number.” But the deeper requirement is that the number has to mean something, achieving the key result must actually advance the objective. The AI industry has inverted this. Benchmark scores have become the objective itself, rather than a key result in service of the real objective: reliable, authorized, verifiable work in production. When the metric detaches from the mission, you get exactly what the benchmark auditing research found, rankings that reshuffle the moment someone checks whether the ground truth was even correct.
Grove also warned about a third trap: unpaired metrics get gamed. He insisted on pairing quantity measures with quality counterparts so that hitting the number could not come at the expense of the outcome. The AI industry is running almost entirely on unpaired metrics, capability scores with no companion measurements for authorization, evidence preservation, reversibility, or real-world validity. Goodhart’s law is doing the rest.
The Orthogonal Take
The AI industry is over-indexing on benchmarks and under-measuring what really matters.
Agents make this problem harder because they turn model outputs into institutional actions. Once an AI system can use tools, modify files, write code, trigger workflows, or interact with credentials, the old measurement frame starts to break. A benchmark score, a task-completion rate, or a successful demo does not prove that the system is reliable in the world.
The better question is not:
How capable is the model?
It is:
What exactly happened, why did it happen, who authorized it, what evidence survived, and would the result hold up outside the benchmark?
That is a Grove-style paired metric for the agent era: capability, paired with verification. Until AI systems can answer that question, and until the industry starts scoring them on it, much of the market will keep confusing motion for progress.


