Seven Signs Your Experimentation Programme Is Theatre
An experimentation programme can look healthy from a distance while producing almost nothing of lasting value. The team is busy. Tests are running. Results are shared in meetings. Dashboards show activity. But when you ask what changed because of all that work, the answer is vague. This is not a competence problem. Most experimentation teams are staffed with talented people. It is a structural problem. The programme is set up to produce experiments, not to produce decisions. Activity is visible. Impact is not tracked. And because nobody is tracking impact, nobody can prove the programme is worth the investment, which makes it vulnerable every budget cycle. Here are seven signs to look for.
1. The team measures itself by test velocity
If the primary success metric is 'number of experiments launched', the programme is optimising for throughput. More tests does not mean more learning. A programme that runs forty experiments a quarter and learns nothing lasting is worth less than one that runs ten and changes how the organisation makes decisions.
Ask: What is our quality rate? Of the experiments we launched last quarter, what percentage had a well-formed hypothesis, a decision protocol, and a documented learning? If nobody tracks this, velocity is the only number available, and it is the wrong one.
2. Nobody can tell you what the programme learned last year
Individual test results exist. But when you ask for the synthesised knowledge, the themes, the patterns, the accumulated understanding of how your customers behave, there is a pause. Someone offers to 'pull something together.' If the learnings have to be reconstructed rather than retrieved, they are not stored in any structured way. They exist in people's memories, old slide decks, and Slack threads. Each departure from the team permanently erases some of that knowledge, and nobody notices because nobody was accessing it anyway.
Ask: Show me the three most important things our experimentation programme has learned in the last twelve months. Time how long the answer takes. If it requires research, the programme is not capturing knowledge. It is losing it.
3. Results are shared but decisions are not tracked
The team presents results regularly. But when you ask what decision was made based on a specific result, the answer is indirect. 'We shared it with the product team.' 'They took it on board.' 'It informed the roadmap.' These are not decisions. A decision is: 'Based on the result of experiment X, we will implement variant B across all markets by the end of Q2. The projected impact is Y. The owner is Z.' If decisions are not documented with that level of specificity, results are being produced and then abandoned. The programme generates evidence. Whether anyone uses that evidence is left to chance.
Ask for the decision log. Not the results log. The decision log. If one does not exist, nobody is tracking the most important output of the programme.
4. Every experiment is a win
A programme that only reports positive results is either not running ambitious experiments or not reporting honestly. Experimentation is about learning, and learning requires testing things that might fail. A win rate above 70% is a warning sign, not a success indicator. It usually means the team is testing safe, incremental changes where the outcome is predictable. These are optimisation tasks, not experiments. They refine what already works rather than discovering anything new. They also tend to produce diminishing returns over time.
Ask: What was our biggest failure this quarter, and what did we learn from it? If the team is uncomfortable with this question, the culture punishes failure rather than learning from it, which means the programme will never test anything genuinely important.
5. Nobody validates whether wins held up after implementation
The team reports a winning test. The variant gets implemented. Everyone moves on. Six months later, nobody has checked whether the projected impact actually showed up in the business numbers. This is the most expensive blind spot in experimentation. It means you are making investment decisions based on test projections that are never verified. It also means implementation quality is not monitored. A variant can be built slightly differently from what was tested, or its impact can be eroded by other changes, and nobody would know.
Ask: Of the experiments we implemented six months ago, how many have been validated? If the answer is zero, every projected impact number you have ever been shown is unverified.
6. The programme cannot survive a team change
If a key person left tomorrow, how much of the programme's knowledge would leave with them? If the answer is 'a lot', the programme is built on people, not systems. That works until it does not, and it stops working at the worst possible moment, when the person with the knowledge is no longer available to answer questions. This is not a criticism of the team. It is a structural observation. If there is no system that captures experiments, results, decisions, and learnings in a searchable, structured way, then the team's collective knowledge is distributed across individual memories and personal files. That knowledge is inaccessible to anyone else and temporary by nature.
Ask: If I needed to understand everything this programme has done in the last two years, and nobody from the current team was available to explain it, where would I go? If there is no clear answer, the programme's history is not recorded. It is remembered.
7. The team asks for more budget but cannot quantify current impact
Budget conversations require evidence. 'We need more headcount because we are running more tests' is not evidence of impact. It is evidence of activity. A programme that can say 'experimentation has directly informed these twelve decisions this year, prevented these three costly mistakes, and validated these two strategic bets before full investment' is making a case that any CFO can evaluate. A programme that can only say 'we ran eighty tests and had a 65% win rate' is not. If your team struggles to make the impact case, it is probably not because the impact is not there. It is because the connection between experiments and business decisions is not being tracked. The evidence exists but it is scattered, undocumented, or stored in a format that nobody can query when the budget conversation arrives.
Ask: If I had to defend this programme's budget to the board next week, what evidence would you give me? Show me, do not tell me.
What to Do About It
If several of these signs are present, the instinct is to question the team. Resist that. The team is almost certainly doing good work within the constraints they have. The problem is usually that the programme lacks the infrastructure to make that work visible, durable, and connected to decisions. Experimentation programmes that are built around execution tooling can run tests efficiently. But running tests is the easy part. The hard part is capturing what was learned, connecting it to decisions, making it searchable, and ensuring it survives personnel changes. That is governance, and most programmes do not have it. Not because the team does not want it, but because nobody funded it. The next time your team presents results, do not ask how many tests they ran. Ask what decisions changed because of the programme. If the answer comes easily, with specifics, your programme is working. If it does not, the gap is not effort. It is structure.