In short: Most AI dashboards count supply: adoption, code, pull requests, agent usage. The measure that matters is validated business value per unit of human intervention: changes that reached production, stayed healthy and moved a declared outcome, divided by the human hours spent on them. In DX's 2026 data, tech-sector AI spend rose about 28-fold while the innovation ratio stayed flat.

What is your AI dashboard actually measuring?

DX's second-quarter 2026 report covers more than 500 engineering organisations, and AI utilisation among them is above 90%. In the tech sector, median quarterly AI spend went from about $1,500 to about $44,000 in a year. Over the same four quarters the innovation ratio, the share of time spent building new features rather than maintaining old ones, barely moved. Developers report saving more than six hours a week, and DX finds little sign of those hours turning into new value across the portfolio.

DX sells engineering measurement, so read these as one vendor's figures. The shape matches what Faros and CircleCI report.

Two lines that used to move together have separated. One is what the organisation spends on AI and how much it uses it. The other is what it gets back. Most board packs still show the first line.

That gap matters more now because of what comes next. The next piece in this series lays out the AI Engineering Production System, the redesign of the whole engineering flow for the agentic era. A redesign needs a target. This piece defines it.

Why do the familiar metrics fail now?

Each familiar metric measures something real. The trouble is where it takes the measurement: at the point of production, which is exactly the part of the system AI inflates. A number taken there rises whether or not the organisation absorbs any of the extra supply.

Faros's 2026 telemetry shows the gap inside a single dataset. Tasks completed per developer rose 34% and epics per developer 66%. Over the same period, bugs per developer rose 54% and incidents per pull request more than tripled. Even a count of finished epics is activity until someone checks what the work did in production.

I have been on the wrong side of this myself. At La Redoute we doubled deployments from 40 to 80 a day, and I led with that number. It was real, and it mattered for flow. It could not tell the board whether the business was changing at the same pace. Deployment frequency is a good flow metric and a weak value metric.

Is coding even where your system is limited?

Every metric in that table starts the clock at the keyboard. For many organisations that is the wrong place to start it. IDC's 2025 study of how developers spend their time put building applications at about 16% of a typical month, with the rest going to requirements, testing, security, pipelines, deployment and operations. Accelerate that 16% as far as you like and most of the system is untouched.

Developer time is itself only part of the picture. Before any intent reaches an engineer, the product system has already decided what to work on, how to split it and who should do it. In the organisations I work with, the constraint often sits there, in one of three recurring patterns.

Too many initiatives in flight. When strategy has not chosen, every team carries a slice of everything and each initiative waits for attention. Agents make it cheaper to start work, which makes this worse.

Initiatives that cannot ship without each other. When planning couples work across teams and systems, each initiative moves at the pace of its slowest dependency. Faster coding in one team only lengthens its wait for the others.

An organisation cut across the flow of value. When team boundaries and decision rights do not match how value flows, every change needs a negotiation, and no coding agent attends that meeting.

These patterns live in the Management and Organisation dimensions of MAMOS, where strategy, portfolio, structure and decision rights sit, and they meet Architecture wherever initiatives need decoupling. For the metric, this widens the system being measured. It runs from strategy to operations, and the engineering function is one part of it.

What exactly is validated business value per unit of human intervention?

Validated business value per unit of human intervention is the value of the changes an engineering system has proven in production, divided by the human hours spent getting them there. It is the signature metric of the AI Engineering Production System, and each of its terms has a precise meaning.

Figure 1. Anatomy of the metric. The numerator counts only change that passes all three gates; the denominator counts every human hour from planning to validation.

Figure 1. Anatomy of the metric. The numerator counts only change that passes all three gates; the denominator counts every human hour from planning to validation.

Validated means proven in use. A change counts once it has reached real users and stayed healthy through a set window, with no rollback and no incident traced to it. Merged, deployed behind a flag or passed in staging is progress. It is still unproven.

I have used 30 days as that window, and it is worth challenging now. Two windows do different jobs. A health window asks whether the change is safe, and with small batches and strong observability it can be shorter than 30 days. An outcome window asks whether the change did what the business wanted, and that takes as long as the business metric needs: a week for a checkout change, a quarter for a pricing one. Keep the two separate. The ratio then moves fast enough to steer by and stays honest enough to report.

Business value means the change moved an outcome declared before it was built. That declaration is the discipline. Each change, or each batch of changes, names the product or business metric it is meant to move. The evidence then has two halves. The quantitative half is the metric itself, such as conversion, cost to serve or time to complete a customer task. The qualitative half is confirmation from users or the business owner that the change did what the intent meant, which a number alone can miss.

There are two ways to count it, and most organisations should start with the first. The simpler option counts validated changes that met their declared outcome; it is easy to audit and hard to argue with. The richer option weights each change by a value class agreed at intent, from routine to strategic, and confirms the class afterwards with the same two kinds of evidence. It says more and is easier to game, so it belongs to the second year.

The idea is older than AI. At La Redoute we deployed changes and validated each one against its requirements through a test plan, an approach we framed as requirements-based engineering. Every requirement arrived with the tests that would prove it, and Cerberus Testing, the open-source test platform that started there, carried those tests from definition through execution to reporting. What has changed is the volume. When agents write most of the code, validation against declared intent is the only definition of done that scales.

Human intervention means every human hour between intent and validation. That covers the planning and alignment that produced the work, writing and clarifying the intent, reviewing, approving, fixing and reworking, rolling back, and responding to incidents the change caused. Planning and alignment hours are shared across initiatives, so allocate them to the initiatives they served; a rough split is better than leaving them out. Planning and intent time count on purpose. Leave them out and the metric rewards pushing work upstream into longer plans and longer requirements, which is where Constraint Migration sends the constraint once coding is relieved.

The denominator is hours rather than headcount or cost for a reason. Human judgment is where the constraint ends up once everything else has been relieved, so a ratio that divides by it measures the one resource agents cannot add.

How does the ratio behave when things go wrong?

Three situations show why the ratio is harder to flatter than anything in the table.

In the first, agents add supply and nothing else changes. Pull requests rise, review hours rise to meet them, and validated change stays roughly where it was. The ratio falls while every activity metric improves.

In the second, review is thinned to keep up. For a quarter the ratio looks better, because intervention hours drop. Then the health window catches the incidents, those changes fail validation, and incident hours join the denominator. The fall arrives one period late. That is the shape of Faros's data, and it is why the window exists.

In the third, the organisation redesigns. Intent is written so agents can act on it, verification scales with supply, and the pipeline can take the merges. Validated change rises while human hours per change fall. This is the only situation in which the ratio climbs, and it is the one an AI programme is supposed to produce.

Can the metric be gamed?

Yes, as every metric can. Three games are predictable, and each has a structural answer.

The first is redefining value after the fact. The answer is to declare the outcome at intent, before build, and count only against that declaration.

The second is hiding human work: review done in private channels, fixes logged as new features. The answer is to measure intervention from systems where possible, through review time, rework commits and incident time, and to treat self-report as the last resort. METR's 2025 trial is the warning here. Experienced developers believed AI had made them about 20% faster, when it had made them 19% slower.

The third is slicing changes thinner to raise the count. Thin slices are good practice. Each slice still needs its own declared outcome and its own review hours, so slicing alone does not raise the ratio.

Alongside all three, report the ratio next to change-failure rate. A ratio that rises while failures rise is a warning, and the board pack should present it as one.

Where does it sit next to DORA and the rest?

This metric does not replace the dashboards engineering leaders already use. It sits on top of them.

Underneath, three layers of diagnostics keep their jobs. Flow and stability metrics of the kind DORA made standard show whether the delivery pipeline works. Adoption and cost metrics, the utilisation and spend that frameworks such as DX's track, show whether the tools are used and what they cost. Stage queues, across the whole product system from planning to operations, show where the constraint sits this quarter. DX draws a similar line between measuring how work is done and measuring whether outcomes improve.

On top sits one economic ratio for the board. The diagnostics explain why it moved. The ratio says whether it moved at all.

Figure 2. The measurement stack. Diagnostics explain the ratio; the ratio is what the board decides on.

Figure 2. The measurement stack. Diagnostics explain the ratio; the ratio is what the board decides on.

DORA's 2025 research points the same way. Its AI Capabilities Model found that a user-centric focus amplifies the benefits of AI, and that adopting AI can harm teams without one. It also found that value stream management strengthens the link between AI adoption and organisational performance. Both findings are, in effect, arguments for measuring validated value.

What goes on the board slide instead?

Replace the adoption slide with four readings.

The ratio, quarter on quarter. Show the trend only, against the same organisation last period, and never as a league table of teams.

Supply against validated change. Draw two lines: what the system produced, and what it proved in production. The distance between them is the Absorption Gap, drawn where a board can see it.

Where the constraint sits, across the whole product system. Name one stage, from strategy and planning through to operations, with the queue in front of it. In some organisations it will be review or the pipeline. In others it will be upstream: initiatives waiting for a decision, a dependency or a team. A board that only ever sees engineering stages will only ever fund engineering fixes.

AI spend per validated change. This is the number that tells a CFO whether a 28-fold rise in spend bought anything.

the replacement board slide. Keep one option, delete the other and its marker before publish.

the replacement board slide. Keep one option, delete the other and its marker before publish.

If you have no baseline yet, the 30-day test in The Absorption Gap produces a first one from data most organisations already hold.

The number to agree before the redesign

The next piece in this series is the redesign itself: the AI Engineering Production System, stage by stage. Every choice in it, from how strategy becomes intent to who owns verification, should be judged by whether it moves this ratio.

If your AI spend doubled next year, what number would tell you it had been worth it?

Read further

FAQ

What is validated business value per unit of human intervention?

It is the value of the changes an engineering system has proven in production, divided by the human hours spent getting them there. A change counts when it has reached real users, stayed healthy through a set window and moved an outcome declared before it was built. Human intervention covers every hour from writing the intent to handling incidents, which makes the ratio a measure of the one resource AI cannot add.

How is it different from DORA metrics?

DORA-style metrics measure how well and how stably the delivery pipeline flows. They can all improve while the business gains nothing new, because they count changes whether or not those changes moved an outcome. Validated business value per unit of human intervention sits above them as the board-level ratio, and the DORA metrics become the diagnostics that explain why it moved.

What counts as human intervention?

Every human hour spent between intent and validation: the planning and alignment that produced the work, writing and clarifying requirements and solution designs, reviewing, approving, fixing, reworking, rolling back and responding to incidents the change caused. Planning and intent time are included on purpose, so the metric cannot improve by moving work upstream. Measure these hours from systems where possible, because self-reported time is unreliable.

Can the metric be gamed?

It can, in three predictable ways: redefining value after the fact, hiding human work and slicing changes thinner. Declaring the outcome before build, measuring intervention from systems and requiring every slice to carry its own outcome close those routes. Reporting the ratio next to change-failure rate catches the remaining case, where review is cut to make the ratio look better.

What should replace the AI adoption slide in a board pack?

Four readings: the trend of validated business value per unit of human intervention; supply against validated change as two lines, which shows the Absorption Gap; the stage where the constraint currently sits, anywhere from planning to operations; and AI spend per validated change. Adoption rates can move to an appendix, since above 90% utilisation they no longer distinguish one organisation from another.

What if AI coding is not our bottleneck at all?

Then faster coding will not move the ratio, and the constraint is likely upstream of engineering. Common causes are too many initiatives in flight, initiatives coupled across teams so each waits for the slowest, and team boundaries that cut across the flow of value. Measuring queues from strategy and planning through to operations, and counting planning hours in the denominator, makes that visible to the board.

Sources