I test landing page copy every day. Here's what works.
At SplitMonk, we run copy experiments constantly -- both for our own site and for customers. After hundreds of tests, I've developed a process that consistently produces winners. It's not complicated, but most teams skip the boring parts and then wonder why their tests come back inconclusive.
Here's the step-by-step guide I actually follow.
Step 1: Pick ONE element. Not five.
I cannot stress this enough. The number one mistake I see is tests that change the headline, the subheadline, the CTA, and the hero image all at once. When that test wins (or loses), you have zero idea what drove the result.
Start with the element highest on the page and closest to the conversion action:
- Headline -- test this first, always. It has the biggest effect on bounce rate and everything downstream.
- CTA button -- test this second. It directly controls the click.
- Subheadline -- test this third.
- Social proof -- test positioning and phrasing.
- Everything else -- feature descriptions, FAQ, etc.
I've tested hundreds of elements across dozens of pages. Headlines produce a median lift of 8-15% when you find a winner. CTA changes produce 5-10%. Button color changes? Usually 0-2%, and rarely significant. Stop testing button colors.
Step 2: Write 3-5 variants. No more.
You want enough variants to explore meaningfully different approaches, but not so many that you need a million visitors to reach significance. I stick to 3-5.
Here's a real example from a SaaS landing page we tested last month:
| Variant | Headline | |---------|----------| | Control | "Build better products with customer feedback" | | B | "Your users are telling you what to build. Start listening." | | C | "Ship features your customers actually want" | | D | "3,200 product teams use [Product] to cut churn by 34%" |
Each variant tests a different angle:
- Control: Generic benefit statement
- B: Pain point + imperative
- C: Outcome-focused, active voice
- D: Social proof + specific number
Variant D won by 17%. Specific numbers almost always beat vague benefits. I've seen this pattern so many times I could write a rule: if your copy doesn't contain a number, add one.
For CTAs, here's the set I test most often:
- "Start Free Trial" (standard)
- "See It In Action" (lower commitment)
- "Get Instant Access" (speed + exclusivity)
- "Start Free -- No Credit Card" (objection removal)
- "Try It Free for 14 Days" (specific + risk reversal)
Step 3: Set your traffic split
50/50 is the default. It's usually right. But not always.
Use 50/50 when: You're early in testing and have no strong prior. You want the fastest possible result.
Use 80/20 (control heavy) when: The variant is risky or untested. You want to limit exposure to a potentially bad change. In SplitMonk, we call this "conservative mode."
Use 33/33/33 when: You're testing 3 variants simultaneously and they're all reasonable bets.
The math: a 50/50 split reaches significance about 2x faster than a 90/10 split for the same traffic volume. If you have limited traffic, 50/50 is almost always the right call.
Step 4: Wait. Seriously, just wait.
Here's what a typical test timeline looks like:
- Day 1: 200 visitors per variant. Data is pure noise. Don't look at the dashboard.
- Day 3: 800 visitors per variant. You'll see big swings -- one variant might show +25%. This is still noise. Do not stop the test.
- Day 7: 2,500 visitors per variant. Results are stabilizing. You might see a clear leader emerging, but wait.
- Day 14: 5,000+ visitors per variant. If you have a winner with >95% confidence and the result has been stable for the last 3-4 days, you can call it.
I have a personal rule: never call a test before day 7, no matter what the numbers say. Weekday vs. weekend traffic has different behavior. Monday-shoppers and Saturday-shoppers respond to different copy. You need at least one full business cycle.
The hardest part of A/B testing isn't the statistics. It's not touching the "stop" button on day 3 when the numbers look amazing.
Step 5: Read the results. Here's what to look for.
When your test completes, don't just look at "winner" or "loser." Here's what I check, in order:
Check the SRM first
Did traffic actually split correctly? If you expected 50/50 and got 55/45, something is broken. Throw out the results.
Look at the confidence interval, not just the point estimate
A result of "+12% lift, 95% CI [2%, 22%]" means: the variant is almost certainly better, but the true lift could be anywhere from 2% to 22%. That's a wide range. If you need at least 10% lift to justify implementation cost, this test doesn't tell you enough.
A result of "+8% lift, 95% CI [5%, 11%]" is tighter and more actionable, even though the point estimate is lower.
Segment by device
I've seen tests where a headline wins by +15% on desktop and loses by -8% on mobile. The aggregate looks like a modest +4% win, but you're actually hurting half your traffic. Always check mobile vs. desktop.
Check your guardrail metrics
A headline that overpromises will boost signups and destroy retention. I always monitor at least one downstream metric -- activation rate, time to first value, or 7-day retention -- alongside the primary conversion metric.
The copy testing stack I recommend
- SplitMonk for variant injection and statistics (obviously biased here, but we built it to solve exactly this problem)
- A spreadsheet for tracking hypotheses, results, and learnings -- I use a simple Google Sheet with columns for element, hypothesis, variants, result, and takeaway
- Customer interviews for generating variant ideas -- the words your users use to describe their problems are almost always better copy than what you'll brainstorm in a meeting
That last point is the most underrated advice in this entire post. Go read your support tickets, sales call transcripts, and G2 reviews. Copy the exact phrases your customers use. Test those against your marketing team's copy. The customer language wins more often than not.

Michał Pogoda-Rosikoń
Founder
Founder of SplitMonk and bards.ai. Data scientist from Wroclaw University of Technology, specializing in NLP and machine learning. Building AI-powered tools that optimize conversions on autopilot.



