← All posts

We Tested the Internet's Most Famous CTA Case Study on Our Own Traffic

Title card: Testing the Internet's Most Famous CTA Case Study on Our Own Traffic

There's a conversion case study you have almost certainly seen if you've spent any time reading about landing pages. It's the one about button copy: change your call-to-action from second person to first person, "Start your free trial" to "Start my free trial," and conversions jump by roughly 90%.

It's cited everywhere. It has been repeated in conference talks, onboarding decks, and about ten thousand listicles, usually with the same screenshot. It is, as far as I can tell, the single most famous button in marketing.

So when we were putting together a batch of landing page experiments on a SaaS product I work on, this one was irresistible. It's the cheapest possible test to build: one string swapped for another. If the famous number held even a tenth of its promise, it would be the best return-on-effort experiment we'd ever run.

We ran our hero CTA copy against a first-person possessive version, fifty-fifty, and let it collect traffic for a couple of weeks alongside our other tests.

It was a wash. No detectable difference in either direction on signups or trials. The most famous lift on the internet, on our page, with our traffic, rounded to zero.

Why famous case studies don't replicate

I don't think the original study was fabricated. I think it was real, in its context, at its time, and that almost everything about how it gets cited strips away the context that made it true. A few mechanisms worth naming, because they apply to every borrowed test idea, not just this one:

1. The baseline does the heavy lifting. A 90% lift says as much about how weak the control was as how strong the variant is. If the original control was a vague, mismatched button on a page with other problems, the variant had room to double. Our control had already survived several rounds of testing. Optimized pages don't have 90% lying around in a pronoun.

2. Context doesn't travel. Button copy interacts with everything around it: the headline, the offer, the audience's mood, the era. First-person copy may genuinely feel more personal in one design system and gimmicky in another. Ripping one element out of a winning page and grafting it onto yours transplants the words, not the win.

3. The original might be noise. Many celebrated case studies come from small samples, tested in an era before sequential-testing discipline was common. Peek at a test early and you'll find dramatic lifts everywhere; most regress to nothing. Some fraction of the internet's classic results are, statistically speaking, coin flips that got screenshotted at the right moment.

4. Survivorship all the way down. Nobody writes a viral post about button copy that did nothing. The case studies you can name are the outliers by construction. For every published 90%, there are unpublished dozens of washes exactly like ours.

5. Time decay. Even a genuinely real effect erodes. Patterns that felt fresh and personal in the year of the original study read as marketing-speak once every SaaS on earth adopts them. Famous tactics carry the seeds of their own irrelevance.

Was the test still worth running?

Yes, and this is the part I actually want to argue for.

The test cost us nearly nothing: one string, one flag, a slice of traffic we were already testing on. In exchange we got a real answer to a question we would otherwise have kept re-litigating in copy reviews forever. Every future debate about first-person button text on this product now ends in one sentence: we tested it, it's a wash, pick whichever reads better.

That's the quiet value of cheap null results. They don't move revenue, but they retire arguments, and retired arguments are how a team's testing roadmap stops being driven by whoever read the most recent listicle.

There's also a discipline benefit. Testing a famous claim calibrates your skepticism in a way no amount of reading can. Once you've watched a canonical 90% lift flatline on your own traffic, you read every future case study differently: not "we should do this" but "here's a hypothesis someone else's traffic once liked."

How to borrow test ideas without borrowing their conclusions

We still mine other people's case studies constantly. The rules that keep it honest:

  1. Treat every case study as a hypothesis, never a result. The result belongs to their page, their audience, their baseline. What's portable is the underlying mechanism.

  2. Prefer the mechanism to the tactic. "Ownership language increases commitment" is a mechanism worth probing in many forms. "Change your button to say my" is one frozen tactic from one page in one year.

  3. Size your expectations by your baseline. If your page has already been through several honest test cycles, expect single-digit effects, and power the test accordingly. Anyone promising you a double from copy alone is describing a broken control.

  4. Budget cheap tests for famous claims. One string swap per testing wave costs almost nothing and either finds free money or permanently closes a debate. Both outcomes pay.

  5. Publish your washes, at least internally. The file of things that did nothing is as valuable as the file of winners. It's the only antidote to the survivorship bias you're consuming everywhere else.

The takeaway

The most famous CTA result on the internet did nothing on our page, and that's the most useful thing it could have done. Not because the original authors were wrong, but because their result was never ours to have. Case studies are other people's answers. Your traffic is the only place your answers live.

Test the famous stuff. Just test it because it's cheap, not because it's famous.