Write the decision rule before you launch
Decide, in the brief, what result causes what action. Both directions. E.g. If the variant lifts trial starts by 5% or more at 95% confidence, Doughnut Labs ships it. Below that, or if average order value drops more than 2%, it reverts. A rule written after seeing data is not a rule, it’s a rationalization, and any flat test can be read as a win once you’re allowed to pick the metric afterwards.Size it before you run it
Work out how long the test needs to run to detect the effect you care about, using statistical significance. If the answer is longer than you’re willing to wait, the honest options are to test a bigger change, accept a lower confidence level and say so, or not run it. Running an underpowered test and reporting the result as if it were conclusive is the most common failure in the whole discipline.Don’t peek
Checking significance repeatedly and stopping when it crosses the threshold inflates false positives substantially. The result looks the same as a real one and isn’t. Peeking covers the mechanism. Watch the guardrail metrics daily for damage. Watch the primary metric at the end.One variable, or you learn nothing
If the variant changes the headline and the image and the audience, a win tells you the bundle beat the control. It doesn’t tell you which part to keep, and the next test starts from scratch. Bundled tests are legitimate when you genuinely want to know whether a whole new approach beats the old one. They’re a mistake when someone wanted three answers and had time for one test.Give it an ID before it exists
Register the experiment ID and apply it to every surface the test touches: campaign and ad names, UTM values, and the analytics events. A variant with no ID cannot be found again once the account has forty campaigns in it, which means the result cannot be reproduced or even located.Record the nulls
A test that showed nothing is a real finding, and the most commonly discarded one. File it in the experiment learning library with the same care as a win. Otherwise the same test gets re-run next quarter by someone who had the same good idea.Roll the result into a standard
A win that stays in a brief is a win that gets forgotten. If a test establishes that something works, write it into the standard it belongs to: the copy brief, the landing page standards, the relevant convention. That’s what turns one result into a default.The failures, ranked by damage
Related resources
- Run an experiment The full path.
- Experiment brief The workbook.
- Creating an experiment The detailed process.
- Statistical significance Sizing and interpreting.
- Experiment review process Judging the result.