Autopilot / Evidence & claims
Evidence & claims
A growth tool is graded on the size of the number it shows you, which is exactly the wrong incentive. Uplift Funnel attaches to every number a label saying how it knows — and enforces the distinction in code rather than in tone.
The four levels
Every figure the product shows came from one of four places, and they are not interchangeable. A dashboard that renders all four the same way is a dashboard that lets an assumption graduate into a result by being printed next to one.
| Level | What it means | Status |
|---|---|---|
| measured_here | Measured on this account, in an experiment with a counterfactual arm. | Available |
| descriptive | Observed counts with no counterfactual. A funnel, not a lift. | Available |
| assumed_prior | A prior from the pattern library. Nothing about this account has tested it. | Available |
| measured_pooled | Measured across accounts on the same archetype — never this account alone. | Not shipped |
The standing rule
A decision taken on a proxy metric may not be reported as a revenue result. measured_here says the effect was measured; it does not say what was measured, and an activation win reported as money is the single most profitable lie a growth tool can tell. Where a reading carries an activation metric, the surface prints the sentence alongside it rather than leaving you to assume.
The metric ladder
When autopilot opens an experiment it does not pick a metric it likes. It picks the highest metric your volume can actually read inside a 28-day window, and it computes that rather than guessing:
| Traffic band | Decision metric | What it may claim |
|---|---|---|
| Small — up to ~100 arrivals a day | High-frequency proxies: step completion, dwell. Descriptive reads on your own funnel. | No lift claim. Priors are labelled as priors. |
| Medium — up to ~1,000 a day | Your own data, on an activation metric: activated, then trial start. | measured_here on activation. Never reported as a revenue result. |
| Large — above ~1,000 a day | Your own data, on money: paid conversion, then revenue net of cost. | measured_here on revenue. Reconciliation becomes readable. |
The ladder climbs on its own: as your traffic grows, new experiments start from a higher rung. A running experiment never changes rung — its horizon is sealed when it opens, because moving the goalposts mid-test is how a result gets manufactured.
The holdout
5% of your users are held back from everything autopilot does, from day one, permanently. It is not a per-experiment control — it is a standing counterfactual for the programme as a whole, which is the only thing that can answer “did all of this add up to anything?”
On a small account the holdout will not be readable for a long time, and the product says so rather than reporting a number from it. It is kept anyway for one reason: when the account grows, the history is already there. A holdout started on the day you need it is a holdout that answers nothing for another year.
Reconciliation
The reconciliation panel puts two numbers side by side: the cumulative lift autopilot has claimed across its promoted experiments, and the lift measuredagainst the holdout. They will not match — winner's curse, novelty effects and interactions all push the first number up — and the size of the gap is the useful part.
Where a lift cannot be measured, it is reported as unmeasurable. Never as zero. Those are different statements and collapsing them would quietly flatter whichever side of the ledger the reader was hoping for.
Where you see this
- Every reading in Results carries its level as a badge.
- The growth memo restates what it did and what it is not entitled to claim.
- The approval gate tells you up front what your traffic will and will not be able to prove — see the approval gate.
- Measurement in the dashboard holds the holdout and reconciliation panels.