Arrio

The measurement problem

How to measure what your software development actually produces

Software is the largest line in the technology budget and the hardest to read. In the AI era it has become harder still. This is a guide to measuring what that investment produces: why the usual numbers cannot tell you, what a real measure answers, and why it has to come from outside.

The scale of the gap

Forecast worldwide IT spending in 2026, up 14.2% on 2025. (Gartner, 2026)
$6.37tn
Forecast worldwide IT spending in 2026, up 14.2% on 2025. · Gartner, 2026
Of CIOs name the assessment of technology ROI as a point of contention with their CFO. (KPMG, 2025)
49%
Of CIOs name the assessment of technology ROI as a point of contention with their CFO. · KPMG, 2025
Of enterprise AI pilots fail to deliver measurable financial return. (MIT, 2025)
95%
Of enterprise AI pilots fail to deliver measurable financial return. · MIT, 2025

The money going into software has never been larger, and the people who sign for it have never been able to see less of what it produces. That gap is the problem this page is about, and it is widening.


Why it is harder now

AI made software faster to produce and far harder to account for.

The old measures, story points, tickets, hours and velocity, were built for a world where effort and output moved together. AI broke that link. A developer can now generate in an afternoon what used to take a week, so counting effort tells you less each quarter about what was actually produced.

The evidence is that more motion is not more value. Across 22,000 developers and 4,000 teams, AI raised task throughput by 33.7% while incidents per pull request rose 242.7% and review times rose 441% (Faros AI, 2026). A controlled study found experienced developers 19% slower with AI on familiar code, even as they felt 24% faster (METR, 2025). A quarter of production code is now AI-authored, and productivity moved about 10% (DX, 2026).

This is the productivity paradox: more is being produced than ever, and leadership can account for less of it than ever.

More incidents per pull request, even as task throughput rose 33.7% with AI. (Faros AI, 2026)
+242.7%
More incidents per pull request, even as task throughput rose 33.7% with AI. · Faros AI, 2026
Slower, in measured time, when experienced developers used AI on familiar code, though they felt 24% faster. (METR, 2025)
19%
Slower, in measured time, when experienced developers used AI on familiar code, though they felt 24% faster. · METR, 2025
A quarter of production code is now AI-authored, and productivity moved about 10%. (DX, 2026)
26.9% / 10%
A quarter of production code is now AI-authored, and productivity moved about 10%. · DX, 2026

Why the usual measures cannot answer it

Activity is not value, and every existing number measures activity.

Engineering dashboards

DORA, SPACE and flow metrics help teams run well and are worth keeping. They describe how work moves, not what it was worth. They report to engineering, not to the budget owner, and they were never meant to.

Self-reporting

Status decks, surveys and steering updates are filtered at every layer on the way up. By the time an answer reaches leadership it has been shaped by the people it assesses. Fewer than a third of leaders who see AI productivity gains can link them to a business outcome (Deloitte, 2025-2026).

Tool vendor telemetry

A tool that measures its own impact has an interest in the result. It answers to the purchase, not to the truth. That is why the measurement has to sit outside the tools, the teams and the vendors it reads.


What a real measure answers

Read the work itself, ground it in cost, report it in business language.

  1. 01

    What did the spend produce?

    Output and quality mapped across every team, division and vendor, on one independent measure, against what each of them costs.

  2. 02

    Is the AI investment paying off?

    The AI-built share of the work, measured against what it delivers and tracked over time, so you can spread what works and stop what does not.

  3. 03

    Is the strategy actually shipping?

    Programmes and projects read against real progress rather than status reports, with course corrections available in weeks, not at year end.

  4. 04

    Where is the risk we cannot see?

    Technical-debt load, security and architecture, read from the code as it stands. AI-induced defects rose 60% even in healthy codebases (CodeScene, 2026), so the unknowns need to be on the table before they price themselves in.

The independent wedge

A measurement is only worth what its independence is worth.

The value of an outside read is that no one in the picture is grading their own work. Independence, in practice, is three commitments: we do not sell the tools we measure, we do not host your code, and we have no incentive to make any number look better than it is. The work is read, not the worker. Access is read-only, sandboxed and audit-trailed.

An estimated 9.5% of engineers contribute little or no meaningful work, an estimated $90 billion wasted every year.

Yegor Denisov-Blanch, StanfordSoftware productivity research, 2024

Questions

The questions worth asking

What does it mean to measure software development?

It means answering, in business terms, what a software investment produces: how much is being built, of what quality, at what cost, and whether it moves the outcomes the organisation is paying for. That is different from measuring activity. Commits, story points, tickets and velocity describe how busy a team is. They cannot tell leadership what the money bought. Measuring software development reads the work itself and grounds it in what the work costs.

Why can we not measure it with the metrics we already have?

DORA, SPACE, velocity and engineering dashboards were built to help teams run well, and they still do. They measure flow and activity. They were designed for a world where effort and output moved together. AI has broken that link: a developer can generate in an afternoon what once took a week, so counting effort no longer tells you what was produced. You need a measure of value, read independently, not a faster way to count tickets.

Is developer productivity even measurable without surveillance?

Yes, because the right unit is the work, not the worker. Arrio reads what an organisation produces at the level of teams, divisions and vendors, from the codebases themselves, read-only and audit-trailed. No keystroke logging, no self-reporting, no ranking of individuals. Leadership gets what it needs to govern the investment, and engineers are measured on the output of the system, not watched.

How is this different from an engineering dashboard?

Dashboards report activity to engineering managers. This measures what the investment produces and reports it to the people who own the budget, in language they can act on. One is an operational tool for running teams. The other is an instrument of governance for the leaders who answer for the spend.

Why does the measurement need to be independent?

Because the parties closest to the work all have a reason to shade the number. A tool vendor wants to show its tool working. A delivery team wants to show delivery. Independence is the wedge: a measure is only trustworthy when the party producing it does not sell the tools it assesses, does not host your code, and has no incentive to inflate the result.

How quickly can an organisation see a first measurement?

In days. A first reading includes history, so the picture arrives with a trend already in it rather than a snapshot that needs a year to mean anything. It starts from read-only access to the repositories in scope, deployed to suit your environment.

Sources

  1. Gartner, 2026 Worldwide IT spending forecast, $6.15 trillion for 2026.
  2. KPMG, 2025 49% of CIOs (against 39% of CFOs) name the assessment of technology ROI as a point of contention; 20% of CFOs satisfied with software visibility.
  3. MIT, 2025 95% of enterprise AI pilots fail to deliver measurable ROI.
  4. METR, 2025 Experienced developers were 19% slower with AI on familiar codebases, while feeling 24% faster.
  5. Faros AI, 2026 Engineering Report 2026 (Acceleration Whiplash), 22,000 developers and 4,000+ teams over two years: task throughput up 33.7% and epics per developer up 66%, against code churn up 861%, incidents per pull request up 242.7%, review time up 441% and bugs per developer up 54%.
  6. DX, 2026 Survey of 121,000 developers across 450+ companies: 26.9% of production code is AI-authored, up from 22% the previous quarter, while productivity gains plateaued at about 10% despite 93% adoption.
  7. Stack Overflow Developer Survey, 2025 84% of developers use AI coding tools; trust in accuracy fell from 43% to 29%.
  8. SonarSource State of Code, 2026 42% of committed code is AI-generated or assisted; 96% do not fully trust AI output.
  9. Deloitte, 2025-2026 79% report productivity gains; fewer than 33% can link them to business outcomes.

See what your software development actually produces.