Most advice about email is somebody's opinion. A/B testing replaces opinion with evidence from your own audience. Instead of arguing about which subject line is better, you send both and let the results decide. This guide explains how to run email tests that actually tell you something, how to avoid the traps that produce fake conclusions, and what to test first when you are just getting started.
What an A/B Test Actually Is
An A/B test is a simple experiment. You create two versions of one element, version A and version B, send each to a portion of your audience, and measure which performs better on a metric you chose in advance.
The whole thing rests on one rule that beginners constantly break: change one thing at a time. If version B has a different subject line and a different image and a different send time, and it wins, you have no idea why. You learned nothing you can repeat. Isolate a single variable so the result points to a single cause.
That discipline is what separates a test from a guess. It is also what makes the result worth acting on next time.
Pick One Variable to Test
You can test almost any part of an email, but some variables move the needle far more than others. Here are the usual candidates, roughly in order of impact.
- Subject line. The biggest lever on whether the email gets opened at all. Our guide to writing email subject lines is a good source of variations worth trying.
- Sender name. "Sam from Greenfield" versus "Greenfield Garden Co." can change open rates more than people expect, and a consistent signature backs up whichever name you choose. You can create a free email signature to keep that identity clean across sends.
- Preheader text. The preview line next to the subject. If you have never written one deliberately, our piece on preheader text is the place to start.
- Call to action. The wording, color, or placement of the button or link you want people to click.
- Send time. Morning versus afternoon, weekday versus weekend.
Start at the top of that list. Subject lines are easy to test, produce clear results, and affect everything downstream, because an email that never gets opened cannot get clicked.
Choose the Metric Before You Send
Decide what "winning" means before the test runs, not after. Picking the metric afterward is how people fool themselves into seeing whatever result they hoped for.
Match the metric to the variable you are testing:
- Testing the subject line, sender name, or preheader? Measure the open rate, because those elements only influence whether someone opens.
- Testing the content, layout, or call to action? Measure the click rate, since the email is already open by then.
- Testing the whole funnel toward a goal? Measure conversions, like purchases or signups, which is the truest signal but needs more volume to read clearly.
A subject line that wins on opens but loses on clicks did not really win. Always keep an eye on the metric one step past the one you are optimizing, so you do not boost a vanity number at the expense of the real goal.
Understand Sample Size in Plain Terms
This is where small lists hit a wall, so it helps to be honest about it. A test result is only trustworthy if you have enough people in each group for the difference to mean something.
Think of it like flipping a coin. Flip ten times and you might get seven heads, which proves nothing. Flip a thousand times and the ratio settles down to the truth. Email is the same: a tiny test produces noisy numbers that look meaningful but are mostly luck.
A few practical rules of thumb:
- If a 2 percent open-rate difference came from a handful of opens on a list of 200, ignore it. It is noise.
- Bigger differences need fewer people to trust. A version that wins 40 to 25 is more believable than one that wins 31 to 29, even at the same list size.
- Most email platforms will flag when a result reaches statistical significance, which is just a way of saying the difference is probably real and not chance. Wait for that flag before declaring a winner.
If your list is small, do not despair. You can still test, but you should test bold changes (two genuinely different subject lines, not two near-identical ones) and accept that subtle tweaks need a bigger audience than you have.
How Long to Run the Test
Give the test enough time to gather a fair read, but not so long that other factors creep in. For most email tests, the bulk of opens and clicks arrive within the first several hours to a day after sending.
A sensible default:
- Send both versions at the same time to comparable, randomly split groups.
- Wait at least a few hours, ideally a full business day, before reading results.
- Avoid running a single test across a weekend if your audience behaves differently then, since that adds a hidden variable.
If you are testing send time specifically, that changes things, because the variable itself is when the email goes out. In that case, our guide on the best time to send email covers how to structure the comparison fairly.
A Step-by-Step Subject Line Test
Here is the whole process on a single realistic example, so you can see how the pieces fit together. Suppose you run a small newsletter for home gardeners and want a better open rate.
Step 1: Form a hypothesis. You suspect a specific, curiosity-driven subject beats a flat descriptive one.
Step 2: Write the two versions. Keep everything else identical.
Version A subject: Greenfield Garden March newsletter
Version B subject: The one vegetable to plant before the frost lifts
Step 3: Set the metric. Open rate, because subject lines only affect opens.
Step 4: Split and send. Your platform sends A to a random half and B to the other half at the same moment.
Step 5: Wait and read. A full day later, version B shows a 38 percent open rate against version A's 26 percent on a list large enough for the platform to mark it significant.
Step 6: Record the learning. You now have evidence that curiosity-driven specifics beat generic labels for this audience. Write that down. Next time, you start from B's style instead of guessing again.
That last step is the one people skip, and it is the most valuable. A test you do not document is a test you have to run again.
Avoid the Traps That Produce False Wins
A few mistakes quietly ruin tests and lead you to confidently do the wrong thing. Watch for these.
- Calling it too early. The first hour often favors whichever version your most engaged subscribers saw first. Wait for the full window.
- Testing on a list too small to read. As covered above, tiny samples produce loud noise. If you cannot reach significance, test bigger changes or pool more sends.
- Changing more than one thing. Worth repeating because it is the most common error. One variable per test, always.
- Running endless tweaks on trivia. Testing the exact shade of a button on a 300-person list is a waste of effort. Spend your tests where the payoff is large.
- Ignoring the result you do not like. If your favorite version loses, the audience told you something. Believe them.
Frequently Asked Questions
What should I test first if I have never done this?
Test your subject line, with two genuinely different approaches rather than two near-identical phrasings. The subject line gates everything else, the result is easy to measure on open rate, and the difference between a great and a mediocre subject is often large enough to read even on a modest list. Once you find a subject style that works, move on to your call to action, then content, then send time.
How small is too small for an A/B test?
There is no hard cutoff, but below a few hundred recipients per version, only big, obvious differences will read clearly, and subtle ones will be drowned in noise. If you have a small list, lean into bold tests and let your email platform's significance indicator be your guide. When it never reaches significance, that is the tool telling you honestly that you do not have enough data to conclude anything, which is more useful than a false answer.
Can I just test send time to boost opens?
You can, and it is a fair variable to test, but treat it carefully. Send both versions to comparable random groups, change only the time, and remember that the best time for your audience may differ from any rule you have read. Many senders find that relevance and a strong subject line matter more than the hour. Test it, but do not expect timing alone to rescue weak content.
Make Testing a Habit
A/B testing is not a one-time project. It is a habit of asking your audience instead of assuming. Change one thing, measure the right metric, wait for a real sample, and write down what you learn. Do that a few times and you will replace a stack of opinions with a short, reliable playbook built from your own results. That playbook compounds, and it is worth far more than anyone else's best-practice list.