Skip to main content
Every completed experiment review gets a row here, linked back to its experiment brief. So does every pilot that reaches its read date. The library exists for one reason: a test whose result nobody can find gets run again. The second run costs the same as the first and produces the same answer, and nobody realises until someone remembers halfway through. Pilots are filed alongside experiments because the question “have we tried this?” doesn’t distinguish between the two. Someone asking whether Google Search works for you needs the pilot’s read whether or not it had a control, and a library holding only controlled tests sends them to repeat it.
Null and negative results are filed with exactly the same care as wins. They are the rows most often skipped and the ones that save the most work, because “we tried that and it did nothing” is only useful if it is written down somewhere findable.

The register

The first three rows are examples using Doughnut Labs, a SaaS company that sells disruptive Doughnut Technology. Delete them and fill in your own.

Status values

Six values, so the register can be filtered rather than read. The first four apply to experiments, the last two to pilots. A pilot never gets Win or Loss. Those words claim a comparison that a pilot has no control to support, and using them here is how a first flight gets cited a year later as proof a channel works. Invalid matters as much as the other three. A test recorded as “no result” when it was actually broken tells the next person the idea doesn’t work, which is not what happened.

Before you design a test, search here

Search on the surface and on the metric, not on the wording of your hypothesis. Someone else will have described the same idea differently. Three outcomes, and all three are useful:
  • Already tested and won. It should already be a default. If it isn’t, that’s the finding.
  • Already tested and lost or showed nothing. Either move on, or state what is different now: more traffic, a different audience, a changed product. “We’ll try it again properly” is not a difference.
  • Not tested. Design it, following run an experiment.

What a row needs to be useful

The register is only as good as the discipline behind the rows.
  • The ID, matching the pattern in experiment ID, so the test can be traced to the campaigns and events that carried it.
  • The hypothesis as it was written before launch, not a version rewritten to match the outcome.
  • The decision rule’s verdict, not an interpretation. Whether it beat the bar is a fact; whether it was worth doing is an opinion, and belongs in the review.
  • What actually changed as a result. A win with no decision recorded is a test that ran for nothing.

When a result stops being true

Results age. A test won in a different market, at a different price, against a different competitive set may not hold. Mark a row superseded rather than deleting it, and link the test that replaced it, so the history of what was believed and when stays intact.
Last modified on August 12, 2026