What Your Team Should Be Telling You But Probably Is Not
Experimentation teams have a reporting problem, and it is not dishonesty. It is self-preservation. The team knows what the programme's real weaknesses are. They know which numbers are inflated, which processes are held together by willpower, and which parts of the operation would collapse if the wrong person left. They do not tell you because the incentives do not reward it. Raising structural problems feels like complaining. Admitting that impact numbers are unverified feels like undermining the programme. Saying 'we need better infrastructure' sounds like asking for budget when the priority is delivering results. So the team reports activity, highlights wins, and quietly manages the mess underneath. Here are the things they are probably not telling you.
We do not actually know our impact
The revenue projections in the quarterly report are calculated from test results and extrapolated across traffic and time. They are mathematically correct and almost certainly overstated. The team knows this. Uplift measured in a controlled test does not always hold in production. Effects decay. Other changes interact. Implementation is rarely pixel-perfect. But nobody goes back to verify. The team does not raise this because the unverified numbers are the programme's best evidence for continued investment. Questioning them feels like undermining the case for the programme's existence.
What to do: Ask for a separate column in the next quarterly report: 'projected impact' and 'validated impact.' The first time you ask, the validated column will be mostly empty. That is not a failure. It is a starting point. The fact that the column exists changes the team's behaviour over time.
Our knowledge walks out the door every time someone leaves
The team knows that most of what the programme has learned is stored in people's heads, not in systems. They have talked about building a knowledge base. They have a ticket for it somewhere. But it never gets prioritised because the pressure to run the next experiment always wins. The team does not raise this because it sounds like an internal problem, not something worth escalating to leadership. But it is a strategic risk. Every departure permanently erases institutional knowledge that cost months or years to accumulate. The next hire starts from scratch, retests things that were already tested, and takes six months to become fully productive.
What to do: Ask the team to estimate, honestly, what percentage of the programme's accumulated knowledge is documented in a way that someone new could access without help. The number will be lower than you expect. That is the measure of the risk.
Quality varies more than we let on
Some experiments are excellent: rigorous hypothesis, solid measurement plan, clear decision protocol. Others are rushed, loosely reasoned, and launched because there was pressure to keep the pipeline moving. The team knows the difference. The report does not show it. Quality governance, when it exists, often depends on one person who reviews plans and pushes back on weak ones. When that person is busy, travelling, or absent, weaker experiments get through. The standard is personal, not systemic. The team does not raise this because admitting inconsistency feels like admitting incompetence, which it is not. It is a process gap.
What to do: Ask for the first-pass approval rate from the review process. This is the percentage of experiment plans that meet quality standards on first submission. If nobody tracks it, quality is not being measured. If it is tracked and varies widely, quality is a function of who happens to be reviewing, not a function of the process.
We cannot prove alignment to strategy
The team believes their work is aligned to business objectives. They can make the case verbally. But if you asked them to pull a report showing every experiment connected to a specific strategic objective, with results and decisions, most teams would need days to assemble it manually. The connection between experiments and strategy exists in people's understanding, not in structured data. That means alignment cannot be audited. It cannot be reported on quickly. And when priorities shift, there is no mechanism to redirect the pipeline because the pipeline is not tagged to objectives in the first place.
What to do: Ask your team to produce, within one hour, a list of every active and completed experiment connected to your top three business objectives this quarter. The time it takes and the completeness of the result tells you exactly how strong the connection is between the programme and the strategy. If it takes days, the connection is reconstructed rather than tracked.
We are running on infrastructure that will not scale
Most experimentation programmes start with spreadsheets, slide decks, and the built-in reporting from whatever A/B testing tool the team uses. These are adequate for a team of three running ten tests a quarter. They are inadequate for anything more ambitious. The team knows this. They know the spreadsheet that tracks experiment history is fragile, incomplete, and only navigable by the person who built it. They know the slide deck archive is unsearchable. They know the testing tool's reporting only covers the test itself, not the decision or learning that followed. But asking for better tooling feels like asking for budget, and the answer has historically been 'make do with what you have.'
What to do: Ask the team to walk you through the tools and systems the programme uses today. Ask them which parts are fragile, dependent on one person, or require manual effort that could be automated. The answer will reveal the infrastructure debt the team has been quietly managing. This is not a wish list. It is a risk register.
The programme's biggest risk is not technical
The team worries about sample sizes, statistical significance, and testing tools. The programme's actual biggest risk is usually political: a senior stakeholder who overrides results when they are inconvenient, a product team that does not use experiment evidence in their planning, or a leadership team that evaluates the programme by volume rather than impact. The team does not raise this because naming a specific senior person as a blocker is career-limiting. But if the culture allows evidence to be overridden by opinion, the programme is producing knowledge that the organisation is not structurally committed to using.
What to do: Ask your team, in a private setting, not a group meeting: 'Has there been a time in the last six months when an experiment result was clear but the organisation did something different anyway? What happened?' The answer will tell you whether your culture actually respects evidence or only respects it when it confirms what leadership already wanted to do.
Why This Matters
These are not complaints. They are structural observations about the gap between how the programme operates and how it reports. The team is not withholding information to be difficult. They are withholding it because the reporting format does not create space for it and the incentive structure does not reward it. If you want the real picture, you have to ask for it directly, in a way that makes it safe to answer honestly. The questions above are designed to surface what the standard quarterly report will never show: whether the programme's infrastructure matches its ambition, whether its claimed impact is real, and whether its knowledge will outlast its current team. Every answer that comes back incomplete or takes longer than it should is evidence of the same underlying issue. The programme is doing good work on top of infrastructure that was never designed to make that work visible, durable, or connected to the decisions it should inform.