Most AI writing tests measure essays, code or stories. Most of what people actually write is shorter: an email to a client, a text to a landlord, a note to a colleague, the message you have been putting off. We built a test for that.
How the test works
CommunicationBench is eighty real requests of the kind people bring to an assistant. Every model answers all eighty. Then every answer goes head to head with another model's answer to the same request, and a judge picks the better one. The pairs are shuffled and run in both orders, so position cannot sway the result, and the judge is a different model from the ones being judged. Scores are on an Elo scale.
It is our own benchmark. We built it and ran it, and we say that up front on the model card.
The results
| Model | CommunicationBench |
|---|---|
| Hemmingway-1 | 1026 |
| Fable 5.1 | 1024 |
| Fable 5 | 1009 |
| GLM-5.3 | 1007 |
| Kimi K3 | 996 |
| GPT-6 Astra | 976 |
| Qwen3.8 27B (base) | 954 |
Grok 4.6 and DeepSeek V4 Pro also came in behind Hemmingway-1. The top two are close: Hemmingway-1 and Fable 5.1 are level within the margin. Hemmingway-1 is fifty points ahead of GPT-6 Astra.
Where the differences are
Averages hide the interesting part.
- Hard asks. The messages you keep rewriting: saying no, asking for money back, telling someone something they will not like. Here the gap is widest. Judged on which reply sounds like a person wrote it, GPT-6 Astra's won 9% of the time and Hemmingway-1's 72%.
- Money and admin, work, talking someone round. Hemmingway-1 wins these, most by a wide margin.
- The wrapper. Fable 5, GLM-5.3 and Kimi K3 bury the message inside commentary, options or notes in more than nine replies out of ten. You then have to dig the email out before you can send it. Hemmingway-1 gives you just the message.
- Where Hemmingway-1 loses: hostile storytelling and long story turns. The story models are better there.
Which to use
For everyday email and messages, Hemmingway-1 and Fable 5.1 lead our test, and Hemmingway-1 is the one that reads as a person wrote it: it also came first in our Human-Likeness test. It is a 27B model with open weights.
In the Hemmingway app it can also read your mail and chats and draft each reply for you. How that works.