What to Measure Instead of Test Velocity
Test velocity is the most commonly reported metric in experimentation. It is also the least useful. Knowing that your team ran thirty experiments last quarter tells you they were busy. It does not tell you whether the programme is healthy, whether it is producing usable knowledge, or whether it is worth the investment. The metrics that actually matter are harder to track, which is why most programmes default to velocity. But if you only measure what is easy to count, you will only optimise for what is easy to count, and your programme will get very good at running a lot of tests that may or may not matter. Here are the metrics worth paying attention to.
Decision Rate
Of the experiments completed in a given period, what percentage had a documented decision recorded? Not 'we shared the results.' A specific decision: implement, iterate, abandon, or investigate further, with an owner and a timeline. This is the single most important metric for an experimentation programme. If experiments are completing without decisions, the programme is generating data that nobody is acting on. A programme with a 90% decision rate and low velocity is outperforming a programme with high velocity and a 30% decision rate every time.
Ask your team to produce this number. If they cannot, decisions are not being tracked. That means nobody can tell you what the programme's actual output is. Not tests. Decisions.
Evidence-to-Decision Connection Rate
Of the decisions that were made, how many can be traced back to a specific experiment, research finding, or data point? And conversely, of the experiments that produced a clear result, how many were connected to a decision that was subsequently executed? This metric reveals whether the programme is influencing the business or operating in parallel to it. A programme can have a high decision rate but low connection if the decisions are being made for other reasons and the experiment results are incidental.
Ask your team: Show me the chain from experiment to decision to implementation for the last five completed experiments. If that chain is not documented anywhere, the connections exist only in someone's memory, which means they are not auditable, not repeatable, and not durable.
Plan Quality Rate
Of the experiment plans submitted in a given period, what percentage passed quality review on first submission? Quality means: falsifiable hypothesis with a 'because' clause, named primary metric with a baseline, documented evidence base, and a decision protocol written before the test launched. This metric tells you whether the team's standards are improving over time. A rising first-pass quality rate means the team is internalising what good looks like. A flat or declining rate means the standards are not embedded in the process. They depend on whoever happens to be reviewing.
Ask your team: What is our first-pass quality rate this quarter versus last quarter? If nobody tracks this, you have no way of knowing whether the programme is getting better at designing experiments or just running more of them.
Learning Capture Rate
Of the experiments completed, what percentage had a structured learning documented that is searchable and accessible to anyone on the team? Not a result. A learning. 'Users respond to social proof more strongly than urgency messaging in checkout' is a learning. 'Variant B won with a 3.2% uplift' is a result. If learnings are not captured in a structured, searchable way, the programme forgets everything it discovers. Each experiment exists in isolation. Patterns go unnoticed. New team members start from scratch. And the same questions get retested because nobody can find the answer from the last time.
Ask your team: If a new hire joined next month and wanted to understand everything we have learned about how our users behave on pricing pages, could they find that in under ten minutes? If the answer involves asking colleagues or searching through old presentations, the programme's knowledge is not stored. It is scattered.
Objective Alignment Rate
Of the experiments running in a given period, what percentage are explicitly connected to a current business objective? Not loosely related. Explicitly linked to a stated goal that leadership cares about. This metric tells you whether the programme is working on things that matter or drifting into self-directed optimisation. Both can produce results. Only one produces results that leadership values.
Ask your team: For each experiment currently in the pipeline, which business objective does it serve? If the answer requires interpretation or mental gymnastics to connect, the alignment is inferred, not structured. That means the programme could drift away from strategic priorities without anyone noticing until the next quarterly review.
Post-Implementation Validation Rate
Of the experiments that won and were implemented, what percentage were checked three to six months later to verify that the expected impact actually materialised in production? Most programmes do not track this at all, which means every projected impact figure in every report you have seen is unverified. This metric is uncomfortable because it sometimes reveals that winning experiments did not hold up. But that is precisely why it matters. If you are investing based on projected returns that nobody validates, you are operating on faith.
Ask your team: Have we ever gone back to check whether a winning experiment actually delivered the impact we projected? If the answer is no, the programme's claimed value has never been verified.
How to Use These Metrics
You do not need all six from day one. Start with decision rate and objective alignment rate. These two alone will tell you whether the programme is producing decisions (not just data) and whether those decisions are connected to things the business cares about. If your team struggles to produce any of these numbers, that is the finding. It means the underlying data is not structured in a way that allows these questions to be answered. The team may be doing excellent work, but if that work is not captured, connected, and queryable, nobody can prove it, defend it, or build on it. That is not a people problem. It is an infrastructure problem. And it is worth solving before the next budget conversation.