Hemmingway Try Hemmingway

Blog · Guide

The best AI for writing emails and messages

We tested the largest AI models on eighty real email and message requests in blind head-to-head matchups. Here are the scores, and what they mean.

By the Hemmingway team ·

Most AI writing tests measure essays, code or stories. Most of what people actually write is shorter: an email to a client, a text to a landlord, a note to a colleague, the message you have been putting off. We built a test for that.

How the test works

CommunicationBench is eighty real requests of the kind people bring to an assistant. Every model answers all eighty. Then every answer goes head to head with another model's answer to the same request, and a judge picks the better one. The pairs are shuffled and run in both orders, so position cannot sway the result, and the judge is a different model from the ones being judged. Scores are on an Elo scale.

It is our own benchmark. We built it and ran it, and we say that up front on the model card.

The results

ModelCommunicationBench
Hemmingway-11026
Fable 5.11024
Fable 51009
GLM-5.31007
Kimi K3996
GPT-6 Astra976
Qwen3.8 27B (base)954

Grok 4.6 and DeepSeek V4 Pro also came in behind Hemmingway-1. The top two are close: Hemmingway-1 and Fable 5.1 are level within the margin. Hemmingway-1 is fifty points ahead of GPT-6 Astra.

Where the differences are

Averages hide the interesting part.

  • Hard asks. The messages you keep rewriting: saying no, asking for money back, telling someone something they will not like. Here the gap is widest. Judged on which reply sounds like a person wrote it, GPT-6 Astra's won 9% of the time and Hemmingway-1's 72%.
  • Money and admin, work, talking someone round. Hemmingway-1 wins these, most by a wide margin.
  • The wrapper. Fable 5, GLM-5.3 and Kimi K3 bury the message inside commentary, options or notes in more than nine replies out of ten. You then have to dig the email out before you can send it. Hemmingway-1 gives you just the message.
  • Where Hemmingway-1 loses: hostile storytelling and long story turns. The story models are better there.

Which to use

For everyday email and messages, Hemmingway-1 and Fable 5.1 lead our test, and Hemmingway-1 is the one that reads as a person wrote it: it also came first in our Human-Likeness test. It is a 27B model with open weights.

In the Hemmingway app it can also read your mail and chats and draft each reply for you. How that works.

Try HemmingwayDownload the app

Read next

  • AI that writes your email repliesHemmingway reads your mail, drafts each reply the way you would write it, and waits for you to send it. Here is how it works and what it will not do on its own.
  • AI that replies to your textsHemmingway drafts replies to your texts, WhatsApp, Slack and Teams messages in the way you write to each person, and only sends on its own when you let it.
  • Hemmingway vs ChatGPT for writingChatGPT does almost everything. Hemmingway does one thing: writing that sounds like you. How they compare on emails, messages, stories and privacy.
  • Hemmingway vs Claude for writingClaude is one of the best writers among AI models. How Hemmingway-1 compares with Fable 5.1, Fable 5 and Opus on emails, sounding human, emotional intelligence and stories.