Why most tests are noise
A test needs enough sends to mean anything. With a few hundred contacts per variant and reply rates in the twenties, a difference of two or three points is within chance. Test one variable at a time, run each variant to at least two or three hundred sends, and only keep differences you would bet on.
The tests worth running
The first line. A specific observation about the company versus a specific observation about the sector. This is the test that moves reply rates most, because it decides whether the reader believes the email was written for them.
The question. A yes/no question versus an open one. Yes/no usually wins on reply rate; open questions sometimes win on the quality of replies. Measure positive replies, not replies.
Length. Eighty words versus a hundred and twenty. Shorter usually wins, but not always: some offers need one more sentence of context.
Send timing. Early morning versus mid-morning, Tuesday to Thursday versus Monday and Friday. Differences are real but small; test this last.
Number of follow-ups. Two versus three. The third often adds replies but also adds the “stop writing to me” replies. Decide based on both.
What not to test
Subject line cleverness, emoji, HTML formatting, images, “P.S.” tricks. These either do not move the number that matters, positive replies, or move it the wrong way.
How to read results
Track replies, positive replies and meetings per variant, not opens. Open tracking is unreliable now that mail clients preload images, and it adds a tracking link that can hurt placement. A variant that gets fewer replies but more meetings is the better variant.
The compounding effect
A single test rarely changes a campaign. Ten tests over a quarter, each kept only when the difference is clear, routinely take a sequence from a fifteen percent reply rate to the high twenties. The method is dull; the result is not.