← All posts

Our Trial Length Test Ran Two Weeks and Told Us Nothing

Title card: Our Trial Length Test Ran Two Weeks and Told Us Nothing

Most experiment write-ups are about winners. This one is about a test that ended with a shrug, because I think the shrug is more instructive than half the wins I've published.

The question was as classic as SaaS questions get: how long should a free trial be? On a product I work on, the trial is short, a few days. The argument for longer is obvious: more risk-free time should mean more people willing to enter a card. The argument against is just as obvious: with a longer trial, urgency evaporates, evaluation drifts, people forget, and the conversion from trial to paid can rot.

Both stories are plausible. Both are told confidently in growth literature. So we tested it: three arms, the current short trial as control against a 7-day and a 14-day variant, with the trial-to-paid conversion as the guardrail.

Two weeks later we ended it with no answer. Not "no difference," which would itself be an answer, but genuinely no answer: the data was consistent with the long arms being better, worse, or the same. We ended it, removed the variant code, and kept the short trial as the default.

Here's the autopsy, because every mistake in it is a mistake I've watched other teams make with much more expensive questions.

Failure #1: three arms, one traffic stream

Multi-arm tests feel efficient: why run two experiments when one can answer everything? The arithmetic disagrees. Splitting traffic three ways instead of two doesn't cut your time-to-answer by a third; it multiplies it, because each pairwise comparison now has a fraction of the sample and the effect sizes you're hunting don't grow to compensate.

We compounded it by weighting the third arm smaller (a defensible caution with the most aggressive variant, and a further tax on power). The result: a test that would have needed many more weeks than our patience or our roadmap allowed.

The rule we took away: arms are expensive, and curiosity is not a reason to add one. If the honest question is "is longer better?", test one longer arm. The ladder (7 vs 14 vs 21) is a second experiment you run only if the first one says longer wins.

Failure #2: the metric was slower than the test

Here's the structural problem with trial-length tests that nobody warns you about: the guardrail metric can't physically resolve faster than the longest trial.

A user who enters the 14-day arm on day one of the experiment cannot convert to paid until day fourteen. For the first two weeks, the arm's trial-to-paid number isn't low, it's undefined. Any dashboard comparing arms during that window is comparing a mature cohort against one that hasn't had its first chance to convert. If you don't gate your readouts on cohort maturity, the long arm looks catastrophic early, then recovers, and the chart whipsaws anyone watching it.

We knew this going in, which is exactly why two weeks was never going to be enough: the 14-day arm needed the test's entire duration just to produce its first fully-baked cohort. A fair readout wanted six to eight weeks minimum. We planned for that abstractly and then didn't protect the runway when other experiments needed the surface.

The general rule: an experiment's minimum duration is set by its slowest metric, not its primary one. If your guardrail takes a billing cycle to mean anything, that's the clock, and you should say so out loud before launch.

Failure #3: we peeked, and the peeks had gravity

With an underpowered test and a slow metric, every interim look is a Rorschach blot. One week in, one arm looked exciting on trial starts. There were conversations. There was a moment where "maybe we just ship it" hovered in the air. That's how underpowered tests do damage even when nobody formally ends them early: they leak noise into decisions through the side door of morale and hallway consensus.

Sequential statistics help, but the deeper fix is procedural: write the earliest legitimate decision date down before launch, and treat every look before it as entertainment.

What ending it actually cost (and bought)

Calling it inconclusive and deleting the variant code felt like admitting defeat. It wasn't. Consider the alternatives:

  • Let it run to completion: six-plus more weeks occupying our highest-traffic surface, blocking a queue of better-powered experiments, to answer a question whose plausible effect had already shrunk on inspection.

  • Ship a winner from the noise: change core billing behavior on a coin flip and inherit whatever the truth turns out to be, plus a false confidence that the question is settled.

Ending it kept the funnel default (which years of revenue already validated in the weakest sense) and freed the surface. The question isn't abandoned; it's queued for a re-run with the lessons applied: two arms, evenly split, sole occupant of its surface, runway pre-committed to the slowest metric's timeline, decision date written down in advance.

What a null result is worth

A genuinely powered null ("we can rule out an effect bigger than X") is a purchasable fact, and often a cheap one relative to the meetings it retires. Ours was weaker than that, an inconclusive rather than a null, but even it bought something real:

  1. It killed the free-lunch narrative. If doubling or quadrupling the trial length produced a lift so large it shone through an underpowered test, we'd have seen it. Whatever effect exists is not enormous. The fantasy of trial length as a growth silver bullet is dead here, and that redirects energy to levers with visible slopes.

  2. It exposed our process debt. The failures above (arm inflation, unprotected runway, undated decisions) were invisible while wins kept arriving. A test that returned nothing made the machinery itself the subject. Every experiment since has launched with a pre-registered decision date and a slowest-metric duration estimate, which is the least glamorous and most valuable thing this test produced.

  3. It taught us to price questions before answering them. Some questions cost two weeks of a side surface. This one costs two months of our most valuable page. Knowing the price list is half of running a real experimentation program; we learned this question's price by failing to pay it.

The takeaway

If you're about to test trial length, or anything whose truth arrives on a billing-cycle delay: count your arms, date your decisions, and let the slowest metric set the calendar. And when a test comes back empty-handed anyway, write it up like it mattered.

Because the write-ups where everything worked? Those teach your readers. The autopsies teach you.