How to Verify a Win After Implementation
There is a category of revenue that exists in experimentation reports and nowhere else. A test wins, the projected annual impact is calculated, the number is reported upward, and the change ships. Nobody ever checks whether the impact appeared. Sum a year of these claims and the programme has, on paper, added more revenue than the company grew. Finance notices this arithmetic eventually, and when they do, the programme's credibility does not recover quickly. Claimed wins that never survive scrutiny are phantom revenue, and the protection against them is a habit most programmes skip entirely: verifying the win after implementation.
Why wins evaporate
A win can be real in the test and absent in production for mundane reasons. The effect regresses toward the mean, because winning tests are selected partly for lucky draws. Novelty fades. The tested audience differs from the full population. The implemented version drifts from the tested version, because what engineering ships six weeks later is rarely pixel-identical to the variant. None of this is misconduct. All of it means the test result is a forecast, not a fact, and forecasts get checked.
The protocol
Verification is decided at decision time, not remembered later. When a winning test is approved for rollout, the approval includes four commitments.
A follow-up window: the date, typically one to three months after full rollout, when the claim is checked. Put it in the calendar at approval, because a verification that depends on someone remembering will not happen.
A holdback where feasible: a small percentage of users kept on the old experience for the follow-up period. This is the gold standard, because it preserves the comparison. Where a holdback is impossible, define the next best check in advance: the metric against forecast, seasonally adjusted, with the comparison method named before anyone knows the answer.
A fidelity check: confirmation that what shipped matches what was tested. This takes an hour and catches a surprising fraction of evaporated wins.
A recorded outcome: the verification result written next to the original claim. Survived, partially survived with a revised figure, or did not survive. The original report and the verification live together, so the repository holds what actually happened rather than what was projected.
When a win does not survive
Treat it as information, not embarrassment. Revise the claimed impact, record why the effect faded if it can be known, and feed the pattern back into how future projections are made. A programme that occasionally reports "this win did not hold, and here is our revised number" is displaying the exact behaviour that makes its other numbers believable.
Why this protects the programme
Verification feels like volunteering for bad news, which is why it is skipped. But the programmes that get defunded are almost never the ones that reported honest, modest, verified impact. They are the ones whose spectacular claimed numbers quietly failed to reconcile with the accounts. Every verified win is a deposit in the only account that keeps a programme alive: leadership's willingness to believe the next number you show them.