A field experiment published July 13 ran 16,880 real work emails through a randomized crossover design: 121 employees, six companies, three weeks. Two of the three conditions were rewritten by GPT-5. Reply rates did not move. Emotional positivity in the finished message did, multiplying the odds of a reply by 3.32.
Outbound teams spent two years auditing whether their sequences sound human. The vendor studies behind that debate disagree with each other by wide margins. Cold email still averages a 3.43% reply rate.
The study covers workplace email, so applying it to cold outreach is an argument rather than a finding. The argument is testable in one sequence this week.
What the experiment actually tested
The paper is titled “Playful AI in Professional Email: A Field Experiment on Tone and Recipient Engagement,” by Ziv Ben-Zion and Teddy Lazebnik. It went up on arXiv as preprint 2607.11749 on July 13, 2026. It has not been peer reviewed, which matters for how much weight you put on it and is worth saying before quoting any number from it.
The crossover design in plain terms
Each of the 121 participants sent their normal work email under three conditions in rotation: writing unaided, having GPT-5 rewrite the message in a playful tone, and having GPT-5 rewrite it in a professional tone.
Because every sender passed through every condition, the design compares a person against themselves rather than comparing one group of people against another. That removes the usual confound in this kind of research, where the teams who volunteer to test AI writing are already different from the teams who do not.
What “within-sender positivity” means and why the distinction matters
The headline result is null. Neither rewriting condition directly changed open rates, reply rates, or response times. The effect showed up one level down, in the variation between a single sender’s own messages.
When the same person sent a warmer message than they usually send, that message performed differently. The condition label predicted nothing. The finished tone predicted a lot.
Why the vendor numbers on AI email contradict each other
Anyone who has tried to settle the AI-versus-human question with published data has run into the same problem. The numbers do not agree, and they do not agree by margins larger than the effect anyone claims to have found.
Three vendor tests, three incompatible gaps
Digital Applied analyzed 100,000 paired sends, matched on persona, ICP firmographics, sequence stage, and sender-domain age, and reported a 4.1% reply rate for AI-generated email against 5.2% for human-written, with spam-flag rates of 8% and 3%.
Saleshandy ran its own test and reported the same 4.1% for AI against 10.4% for humans. Prospectory tested 10,000 emails and landed at 8.2% for AI against 11.7% for humans.
All three are vendor-published. None has been independently audited. The human baseline moves from 5.2% to 11.7% depending on who ran the test, which is a spread more than twice the size of the AI penalty any of them claims to have measured.
What a randomized crossover design controls for that a vendor test does not
A vendor comparing its AI output against a human sample is comparing two different sets of senders, two different lists, and often two different time periods. Randomization and rotation remove all three.
The mechanism: Tone as an indirect pathway
The interesting part of the paper is the shape of the effect rather than its size.
What the editing conditions moved
Playful rewriting raised emotional positivity by a coefficient of +0.068 (p<0.001). Professional rewriting lowered it by 0.041 (p<0.001). Both effects are statistically significant. Both stop there.
The rewriting changed the language and the language changed nothing about recipient behavior on its own.
Reading an odds ratio correctly
Within-sender positivity predicted opening at an odds ratio of 2.05 and replying at 3.32 (p<0.001). An odds ratio of 3.32 does not mean a 3.32-in-1 chance of a reply, and it does not mean replies tripled.
It means the odds of a reply multiplied by 3.32 as positivity increased within a given sender’s messages.
The authors describe this as an indirect pathway: the AI editing shaped behavior only by shaping tone, with no direct effect of its own.
For an outbound team, the practical translation is that the tool you write with is not the variable. The register the message lands in is.
What this study does not prove
This section exists because the transfer from the paper to your sequence is where the reasoning gets thin, and skipping past it would be the same move the vendor studies make.
Workplace email carries a prior relationship that cold outreach does not
Every message in the experiment went between people who already knew each other. Colleagues, clients, existing threads. A warm recipient reading a warmer-than-usual note from a known sender is not the same situation as a stranger reading a first-touch pitch.
The mechanism the study found is defined as variation within a sender’s own established correspondence. Cold outreach has no established correspondence to vary against.
Preprint status, single experiment, no replication
One study, 121 people, six companies, not yet peer reviewed. Treat the direction as credible and the magnitude as provisional.
What a cold-outreach version would have to measure
To close the gap properly, someone would need to run the same rotation on first-touch cold sequences, hold list quality and sending infrastructure constant, and measure spam placement alongside replies. That study does not exist yet. Until it does, tone-as-lever is a hypothesis with good mechanical logic behind it.
A tone pass you can run on one sequence this week
Scoped deliberately small. One sequence, one week, one variable.
The five steps
- Pick a sequence with volume history. You need at least a few hundred sends of baseline before the change so the comparison means something.
- Score the current copy for register, not quality. Read each step and mark it warm, neutral, or clipped. Most outbound sequences skew clipped, because efficiency edits strip warmth first.
- Rewrite one step, not the whole sequence. Raise the warmth of a single email while holding the subject line, offer, CTA, and send timing fixed. Changing several things at once produces a result you cannot attribute.
- Run it for a full cycle and watch replies plus spam placement together. The vendor data disagrees on almost everything except that tone changes and deliverability changes tend to arrive at the same time.
- Compare against the same sequence’s own history. Within-sender comparison is the whole design of the study. Benchmarking against an industry average tells you much less than benchmarking against yourself. If your stack classifies inbound replies by type, split the comparison further: a warmer message that lifts total replies while lifting the share of polite declines has not moved anything you can bank.
Running this when an agent writes the copy
If an AI sales agent is generating your messages, the study’s finding lands closer to home. The agent is already choosing the register, and in most stacks nobody has ever specified it.
AnyBiz agents decide timing, channel, and content per prospect across email, LinkedIn, and phone, which means register is a setting somebody should be reviewing rather than a byproduct of whatever the model defaulted to.
Pull ten sent messages from a live campaign and read them the way a recipient would. If they all read clipped, you have found your variable before running a single test.
Mistakes that make the pass useless
Warmth is not an exclamation point, and it is not opening with a compliment about someone’s recent LinkedIn post.
The study measured emotional positivity in the finished message, which in practice looks like plain language, an acknowledgment that the recipient has other things going on, and a request sized to what a stranger would reasonably do.
Piling on enthusiasm moves a different variable and usually moves it the wrong way.
The reply-to-placement argument, stated as an argument
There is a further step most deliverability content takes without flagging it, so this piece will flag it.
What mailbox providers document and what they do not
The standard claim is that replies are the strongest positive engagement signal for inbox placement. Google and Microsoft publish extensive guidance on authentication, complaint rates, and bulk sender requirements.
Neither publishes documentation that ranks replies against other engagement signals or quantifies their weight. Every source that states the claim confidently is a vendor with a product attached to it, and that includes the warm-up category.
Why the reasoning holds, and where it stops
The mechanical logic is reasonable. A reply is expensive for a recipient to fake; it creates a two-way thread, and it is the hardest signal for a sender to manufacture at scale. Filtering systems that weigh engagement have an obvious reason to weigh it heavily.
That reasoning is sound and it is still inference. Treat reply lift as a plausible second-order deliverability benefit rather than a documented one, and do not build a forecast on it.
What you can act on is the layer underneath, which is documented and boring: authentication, domain warm-up, complaint rates, and sending volume discipline.
AnyBiz plans include warm-up and managed deliverability for this reason, since a tone experiment run on a domain that is already filtering will return noise. Get placement stable first, then test register against it.
What to do next
The finding worth acting on is narrow and testable: the register of the finished message predicted reply behavior in a controlled setting, and the tool that produced the message did not.
That test only pays off if you are measuring the right thing on the other end.
An outbound program is worth judging on pipeline and cost per meeting rather than on send volume, and a tone change that lifts replies without lifting booked meetings is a result you want to catch early. Set the measurement up before you run the experiment.
Get your baseline before you rewrite anything. Start with AnyBiz and see what your current sequence is producing today.
