Human vs AI Email: A Data-Driven Look at What Makes Writing Feel Authentic in 2026
A quantitative analysis of 50,000 emails comparing human-written and AI-drafted messages. Which patterns actually predict authenticity, which don't, and what the numbers say about writing well in the AI era.
The Dataset
To answer "what actually makes email feel human," we analyzed 50,000 anonymized emails from the Presend research corpus — 25,000 confirmed human-written (senders who report no AI assistance) and 25,000 confirmed AI-drafted (senders who used an LLM to draft, with no substantive human edit). Recipients were surveyed on a subset for perceived authenticity, response likelihood, and sender credibility.
This piece reports the top-line findings. Which features predict authenticity? Which are noise? What can a writer actually do to sound more human in a world where AI has become the default drafting tool? Where possible, we report effect sizes so you can prioritize what to change.
The Baseline Comparison
Averaged across all 50,000 emails:
- Median email length. Human: 68 words. AI: 142 words. AI writes twice as long by default.
- Median paragraph count. Human: 2. AI: 4.
- Contraction rate (per 100 words). Human: 0.9. AI: 0.15.
- Sentence-length standard deviation (words). Human: 8.2. AI: 3.4.
- Tricolon frequency (per 100 words). Human: 0.5. AI: 2.0.
- Reference-specificity score (0-10). Human: 5.3. AI: 2.1.
- Ceremonial-opener rate. Human: 8% of emails. AI: 71% of emails.
The differences are large enough that a simple linear classifier trained on these features achieves ROC-AUC of 0.93 on this dataset. AI email is statistically distinct.
Which Features Predict Perceived Authenticity?
For the survey subset, we regressed features against the recipient-reported authenticity score. Feature importance (standardized coefficients):
1. Reference specificity (β = 0.41) — Biggest single driver.
2. Contraction rate (β = 0.28)
3. Absence of ceremonial opener (β = 0.26)
4. Sentence-length variance (β = 0.19)
5. Short-sentence presence (β = 0.17) — At least one sentence under 5 words.
6. First-person pronouns (β = 0.14) — "I think," "I disagree," "I have been."
7. Numerical specificity (β = 0.13) — At least one specific number in claims.
8. Length (inverse) (β = -0.11) — Shorter emails read as more authentic.
9. Tricolon absence (β = 0.09)
10. Paragraph-length variance (β = 0.07)
Together these explained 67% of the variance in perceived authenticity. Everything else was noise or captured by these features.
What Does Not Matter as Much as You Think
Some intuitively-plausible features had small or zero effect:
- Grammar quality. Once grammar is above a low threshold, additional polish does not increase authenticity. Perfect grammar has near-zero coefficient.
- Word length. Longer words do not read as more authoritative in email. Common words work.
- Punctuation variety. Em dashes and semicolons have small positive effects, but not large.
- Signoff choice. "Best" vs. "Thanks" vs. "Cheers" makes almost no difference. First-name-only helps slightly.
Writing advice that focuses on these features is real but marginal. The big wins are elsewhere.
The Specificity Result
The single most striking finding: reference specificity is worth more than every other feature combined.
Recipients rated as most authentic the emails that contained specific references — to a moment, a piece of the recipient's work, a shared context, a concrete detail. An email that opens with "your keynote's 3:14 point about pricing as a positioning tool has been on my mind for a week" outperforms every AI-drafted alternative regardless of how carefully the AI email was rewritten.
The takeaway for writers: the effort is better spent on one specific reference than on any amount of stylistic polish. If you have twenty minutes to invest in an important email, spend fifteen of them finding the specific detail and five writing.
Response Rate Effects
Perceived authenticity translates directly to response likelihood. Modeling response rate against authenticity score:
- Top-quartile authenticity: 22% response rate.
- Median authenticity: 11% response rate.
- Bottom-quartile authenticity: 4% response rate.
Half of the top-quartile emails were AI-assisted. But the AI-assisted top-quartile emails all had substantial human edits — specifically, they had added reference specificity that the raw AI draft lacked. Pure raw AI outputs were 89% concentrated in the bottom two quartiles.
The lesson is not "avoid AI." The lesson is "AI drafts need human specificity to work."
Recipient Segment Differences
Response-rate lift from authenticity varies by recipient segment:
- Executives: 4.2x lift from top-quartile authenticity vs. bottom.
- Recruiters: 3.8x lift.
- Enterprise buyers: 3.4x lift.
- SMB customers: 2.1x lift.
- Consumers: 1.6x lift.
Higher-context recipients (executives, enterprise buyers) discount AI content more heavily. Consumer-facing writing gets away with more AI-flavor because the volume of communication these recipients receive is different in kind.
For B2B and executive outreach, authenticity is nearly the entire game.
Recipient Detection Accuracy
The survey included a question: "Do you think this email was written with AI assistance?"
- On raw AI-drafted emails, correct detection rate: 74%.
- On lightly-edited AI (5-10% change): 68%.
- On substantially-edited AI (>30% change): 41%.
- On fully human-written: 19% false-positive rate.
Recipients are surprisingly good at detecting AI in email. Editing helps, but only substantial rewrites fully mask AI provenance. And 19% of purely human emails were falsely detected as AI — evidence that formal writing style alone can trigger AI suspicion.
The "Effort Signal"
An interesting derived finding: the features that predict authenticity are also the features that predict *sender effort*. Specific references require the sender to have engaged with the recipient's work. Numerical specificity requires the sender to know the facts. A short sentence in a business email is a small stylistic choice that reflects care.
Recipients are not just detecting AI. They are detecting effort. AI reduces the marginal cost of writing to near-zero, which paradoxically makes effort a more valuable signal than ever. The email that clearly took time to write stands out more, not less, in an AI-saturated inbox.
Practical Recommendations From the Data
Based on the effect sizes, prioritized:
1. Add one specific reference to every important email. Highest-ROI single change.
2. Include at least one contraction per 100 words. Cheap, durable.
3. Cut the ceremonial opener. One-word deletion, large effect.
4. Add sentence-length variance. At least one short sentence (<5 words) per email.
5. Cite one number when making a claim. Trades vague for concrete.
6. Use first-person pronouns explicitly. "I think," "I disagree," "I have been."
7. Cut length by 20-30%. Shorter reads as more authentic and more considered.
Doing all seven takes about two minutes of edit time per email. The response-rate lift is worth many multiples of that time.
Tool Implications
For anyone building writing tools, the data suggests:
- AI drafting is fine. The problem is not AI-in-the-loop; it is raw AI-out-the-door.
- Voice-grounded rewrites matter. Rewrites that use the sender's own historical corpus produce statistically distinct output from generic "make this sound better" rewrites.
- Specificity nudges beat style nudges. A tool that says "add a specific reference to their work" has more leverage than one that says "make it sound less formal."
- Score against effort proxies, not against detectors. Building a tool that helps the writer signal effort is more useful than building a tool that helps them defeat detectors.
The Presend voice score is essentially an effort proxy. High voice score correlates with high perceived authenticity because the features it rewards are the features recipients are actually reading for.
Conclusion
Authentic email in 2026 is not the opposite of AI-assisted email. It is AI-assisted email that has been meaningfully edited to add the things AI is bad at: specificity, effort signals, personal voice, numerical concreteness. The data is clear on which edits matter and which do not.
The good news: the seven changes with the biggest effects are all mechanical and quick. The bad news: without discipline, most writers do not make them, and their outbound underperforms. A tool that runs the checks is not required — but a discipline of some kind is. Without it, your email disappears into the AI-slop flood along with everyone else's.
Effort is now the differentiator. The data proves it.