Skip to main content
This page is a workbook. Fill in sections A to E before the pilot launches, and section F on the read date. Each section has a concept review dropdown explaining what it is for. Grayed E.g. text is a placeholder. Delete it and replace it with your own. Examples use Doughnut Labs, a SaaS company that sells disruptive Doughnut Technology.
Use this when you are spending to find out whether something works and there is no control to compare against. If you have a control, you are running an experiment: use the experiment brief and creating an experiment instead.

Pilot or experiment?

A pilot cannot produce the thing an experiment produces. There is no control, so there is no lift, no significance, and no defensible claim that the result was caused by the change rather than by the month it ran in. Treating a pilot as an experiment invites exactly that overclaim, which is why the experiment process turns it away at the door.What a pilot can produce is a decision made against a threshold nobody could move afterwards. That is a weaker form of evidence, and it is still worth a great deal more than the usual alternative, which is a channel that quietly continues because nobody remembers what it was supposed to prove.The two records share a library for a practical reason. The question “did we already try this?” does not distinguish between a controlled test and a first flight, and an answer that only covers one of them sends someone to repeat the other.

A. Setup

Pilot ID source: E.g. Issued from the same log as experiment IDs, using the PILOT prefix so the two can’t collide. See experiment ID convention.
The pilot ID does the same job as an experiment ID: it is the string that ties this record to the campaigns, the rows in the flat file, and the eventual decision. Without one, the pilot’s spend is indistinguishable from ordinary spend three months later, and the question of what it cost gets answered by reconstruction.Issuing it from the same log as experiment IDs, with a different prefix, avoids the failure where two systems independently generate 04. It also means one search finds every deliberate piece of learning spend, whichever form it took.The read date is filled in here, at setup, rather than left to be picked later. A read date chosen after seeing the numbers is not a read date.

B. The question

What are you buying information about, and what would change depending on the answer?
The last two rows are what separate a pilot from spending money and seeing what happens. A pilot whose “no” answer leads to “run it another quarter and see” was never a pilot, and the budget would have been better spent on the thing already working.Writing the “yes” action down in advance also sizes the pilot correctly. If a positive answer would move a tenth of the budget, a pilot that costs a third of the budget has already lost, whatever it finds. That check is only available before launch.The “what we are not asking” row is there because pilots attract scope. Every question that would be interesting to answer arrives during setup, and each one added is another reason the read will be inconclusive.

C. Budget and duration

Check the budget against the volume the question needs before launching, not after. A pilot that ends with “too few conversions to tell” has spent its whole budget buying nothing. The platform brief for your channel carries the volume thresholds: see budget sufficiency on Google or the equivalent section in the TikTok, Meta, or LinkedIn brief.
The sufficiency row is the one that most often turns a planned pilot into a different plan. A budget picked because it is what was available, rather than because it buys enough events to read, produces a number with an error bar wider than the decision it is meant to inform. Working backwards from the volume the question needs is the same arithmetic budget planning applies to a campaign target, run against information rather than results.The ceiling exists because pilots drift. Spend rises to keep a promising signal alive, the read date moves with it, and eighteen months later the pilot is a channel nobody decided to run. A stated hard stop is what makes the read date real.Separating media from production matters more on a pilot than on a campaign, because production is a larger share of a small budget. A $15,000 pilot that spends $5,000 on a landing page is a $10,000 pilot, and judging it on $15,000 of spend understates it.

D. What you’ll measure

How this pilot is identified in reporting: E.g. Pilot ID in the campaign name and in utm_campaign, so its rows separate cleanly in the flat file.
A pilot has no control, so every metric here is being read against something else: a historical baseline, another channel, or a target derived from unit economics. Naming which of the three, per metric, is what stops the read turning into an argument about whether $45 is good.The guardrail row is the one most often left out and the one that most often matters. A new channel can deliver cheap signups that convert to paid at half the rate, and a pilot measuring only cost per signup will read that as a win and scale it.Definitions link out rather than getting restated. A pilot that invents its own version of cost per lead produces a number that cannot be compared to the channel it is being judged against, which was the entire point of measuring it.

E. The decision rule

Write the rule before launch. On the read date you apply it; you do not revisit it. Who decides: E.g. Head of growth, on the read date, against this table. Number of extensions allowed: E.g. One, and only into the “continue” branch above.
A “continue” outcome with no second read date and no extension limit is how a pilot becomes permanent without a decision. Fill in both.
This table is the whole point of the page. Everything above it is setup; this is the part that has to exist before the numbers do, because a threshold written after seeing the data is not a threshold, it is a description.Three outcomes tends to work better than two. A pure pass or fail forces a promising but unclear result into one bucket or the other, and the honest answer for a first flight is often “the signal is real but the read is not clean yet.” Naming that as its own branch, with a limit on how many times it can be taken, keeps it from becoming the default.Naming the decider matters as much as naming the condition. A rule with no owner gets applied by whoever is most invested in the answer.

F. The read

Fill this in on the read date, not before.
Recording the outcome against the rule, rather than describing what happened, is what makes a pilot findable later. Someone asking “have we tried Google Search?” needs the outcome and the threshold it was judged against, not a paragraph.The two open rows at the bottom carry most of the value that survives a year. A pilot that hits its number still teaches you what the number hid, and a pilot that misses is frequently more useful than one that succeeds, provided somebody wrote down why. Both belong in the library, in the same place experiment results go, because the person searching does not know in advance which form the answer took.

Filing

1

Register the ID before launch

Issue the pilot ID and add a row to your experiment log, so the ID cannot be reused. See experiment ID convention.
2

Apply the ID to every surface

Campaign names, utm_campaign, and any dashboard filter, so the pilot’s rows separate from ordinary spend in the flat file.
3

File the brief when it is written

Into the campaign library, named per the file naming conventions.
4

File the read into the learning library

Once section F is complete, add it to the experiment learning library with its outcome, so the next person asking finds it.
Last modified on August 10, 2026