Release-preview draft: this guide describes the planned confidence workflow. Availability and final labels must be confirmed before this guide is published as live product documentation.
Start with one success metric
The planned first version supports CVR, AOV, and RPS as success metrics.
Choose the success metric before evaluating results. Engagement measures such as scroll depth can help explain behavior, but are not included in this first set of confidence calculations.
Lift: how much the result differs
Lift is the percentage difference between a variant’s selected metric and the control’s metric. For example, moving from 4% CVR to 4.4% CVR is +10% lift. It is also a 0.4 percentage-point increase in CVR; these are two ways to describe the same change, not interchangeable numbers. A negative lift means the variant’s observed result is lower. The control is the baseline. When the control metric is zero, relative percentage lift is not defined and should display a dash rather than an infinite improvement. Lift describes what has happened so far. It does not guarantee that the same improvement will continue.Confidence: how much evidence supports the difference
In the planned Bayesian workflow, confidence estimates how likely the variant is to outperform the control on the selected success metric, given the data and the statistical model. For example, 92% confidence means the model assigns a 92% probability to the variant outperforming the control on that metric. It does not mean 92% of visitors will buy, the result is 92% better, or there is a 92% chance of reproducing the exact observed lift. A large lift with little data can still be uncertain. A modest difference supported by more consistent evidence can have higher confidence.Set a confidence threshold
The planned configuration offers 85%, 90%, 95%, or a custom threshold. The threshold determines when a qualifying positive result can be shown as Winning. A higher threshold requires stronger evidence; it does not guarantee a bigger commercial improvement. Keep the threshold aligned with your decision before reviewing outcomes. Lowering it only because a preferred result has not reached it makes the decision less meaningful.Read the result status
The control remains the reference baseline. Minimum-data requirements take precedence over a strong-looking confidence number. Treat the requirements shown by the released product as the source for the applicable threshold; a universal session count is not proof of a reliable result.
