How to Read an Experiment Report Without Being Misled
Experimentation teams present results. Leaders make investment decisions based on those results. The problem is that most results presentations are designed to demonstrate activity, not to inform decisions. They show what was tested and what the data said. They rarely show what decisions were made, what was learned, or whether the claimed impact actually materialised. This is not deception. It is a format problem. The team reports in the format they know: test, result, statistical significance. But that format leaves out everything a leader actually needs to evaluate. Here is what to look for and what to question.
Watch for Phantom Revenue
The most common inflation in experimentation reporting is revenue attribution. A team runs a test that increases conversion rate by 2% on a checkout page. They multiply that by annual traffic and revenue per transaction and report a projected annual impact of several million. The number is technically calculated correctly. It is also almost certainly wrong. It assumes the effect persists permanently at the same magnitude. It assumes nothing else changes. It assumes the winning variant was implemented correctly and completely. It assumes the measurement was accurate in the first place. In practice, most projected impacts shrink significantly once the variant is in production, and very few teams go back to verify.
When you see a revenue projection, ask: Have we validated this? Did we go back and check whether the projected impact actually showed up in the business numbers three months after implementation? If nobody has checked, the number is a forecast, not a fact. Treat it accordingly.
Distinguish Results from Decisions
A result tells you what happened in a test. A decision tells you what the organisation did about it. These are different things, and a report that only contains results is reporting activity, not impact. For every result presented, you should be able to see: what decision was made, who made it, when it was made, and whether it has been implemented. If the report says 'the test won with 95% statistical significance' but does not say what happened next, nobody is tracking the most important part.
For each result in the report, ask: What did we do with this? Where is the decision documented? If the answer is 'we shared it with the product team', that is not a decision. That is a handoff with no accountability for what happens next.
Check for Cherry-Picking
Not all cherry-picking is deliberate. Sometimes the team genuinely believes the programme is in good shape because they are presenting their best work. But a report that only shows wins is not a performance report. It is a highlight reel. Look for balance. A healthy programme produces wins, losses, and inconclusive results. If everything in the report is positive, either the team is only sharing good news, or they are not running experiments that are ambitious enough to fail. Both are problems. Also watch for metric selection. A test might have failed on the primary metric but shown a positive movement on a secondary metric. If the report leads with the secondary win, the framing is misleading even if the data is accurate.
Ask: Of the experiments completed this quarter, how many won, how many lost, and how many were inconclusive? What did we learn from the losses? If nobody can answer the second question, negative results are being discarded instead of mined for knowledge.
Question the Baseline
When a team reports a percentage improvement, the baseline matters enormously. A 15% improvement sounds impressive. A 15% improvement on a page that receives 200 visits a month is not. Similarly, watch for experiments that test against a weak control. If the current experience is obviously broken and the variant fixes it, the 'win' is not evidence of good experimentation. It is evidence that someone noticed a problem and fixed it, which could have been done without an experiment at all.
When the improvement sounds impressive, ask: What is that as an absolute number? How many users were affected? And was the control the best version we had, or was it already known to be underperforming?
Look for the Learning, Not Just the Outcome
The most valuable experiments are often the ones that lost. A clear loss tells you something definitive about how your users behave. That knowledge prevents future mistakes and redirects effort toward approaches that are more likely to work. If the report treats losses as failures rather than learnings, the programme is optimising for a win rate rather than for knowledge. A programme with a 90% win rate is probably not testing anything interesting. The most innovative experiments have high failure rates because they are testing genuine unknowns.
Ask: What was the most surprising result this quarter, and what did it change about how we think about the product? If nobody can name one, the programme is confirming assumptions rather than challenging them.
Verify the Implementation
A test result is only as valuable as the implementation that follows it. Many winning experiments lose their impact during implementation because the variant is not built exactly as tested, or because it interacts with changes made by other teams, or because it was only partially rolled out. Very few experimentation programmes track post-implementation performance. The test wins. It gets handed off. Everyone moves on. Whether the projected value actually materialised is nobody's job to check.
Ask: Of the experiments we implemented last quarter, how many have been validated to confirm the expected impact showed up in production? If the answer is none, you are making investment decisions based on projections that nobody verifies.
What a Good Report Looks Like
A report you can trust is one that shows results, decisions, and learnings together. It includes wins and losses. It states projected impact and separately states validated impact. It connects results to business objectives so you can see whether the programme is working on things that matter. And it tells you what the programme needs from you, not just what it produced. If the report you currently receive does not do these things, the team may not have the systems in place to produce them. Connecting experiments to decisions, tracking implementation, validating impact, and synthesising learnings across experiments requires more than a spreadsheet and good intentions. It requires infrastructure that most programmes do not have, which is why most reports default to the easy version: here is what we tested, here is what won.