Anthropic's Claude models have a reputation as good writers, and in our tests they earn it. Fable 5.1 was the closest model to Hemmingway-1 on everyday writing, and Fable 5 Max wrote better long stories. So this is a close comparison, and we will say where each wins.
The numbers
| Test | Hemmingway-1 | Claude |
|---|---|---|
| Everyday messages (our test) | 1026 | Fable 5.1 1024, Fable 5 1009 |
| Sounds like a person (our test) | 1032 | Fable 5.1 1006, Fable 5 992 |
| EQ-Bench 4 (public) | 1330 | Fable 5 1341, Opus 4.7 1312, Opus 4.8 1285 |
| Long stories (our test) | 1197 | Fable 5 Max 1277 |
Where Claude wins
- Long stories. Fable 5 Max beat Hemmingway-1 on StoryBench by 80 points, and the newer Fable 5.1 is stronger still. For a novel or long chapters, Claude is the better pick.
- EQ-Bench 4. Fable 5 scored eleven points above Hemmingway-1 and was first in the comparison.
- Everything that is not writing. Claude is a general assistant: code, analysis, long documents. Hemmingway is not.
Where Hemmingway wins
- Sounding like a person. Asked which of two replies a person wrote, the judge picked Hemmingway-1's over Fable 5.1's more often: 1032 to 1006.
- Giving you just the message. Fable 5 put the message inside commentary, options or notes in more than nine replies out of ten in our test, and Fable 5.1 did it often too. Hemmingway-1 gives you the text you asked for and nothing around it.
- Everyday messages. Level with Fable 5.1, ahead of Fable 5.
- Size and openness. Hemmingway-1 is a 27B model with open weights under Apache-2.0. You can run it yourself. Details here.
How the app is different
Claude is a chat assistant. Hemmingway is built around the people you write to. You tell it who someone is and it writes to them the way you would. In the apps for Mac, Windows and Android it reads your mail and chats and drafts each reply for you to send, and Autopilot can answer the easy ones by your rules.
We never train our models on your chats, memory or People. Our privacy page lists everything we keep.
About the tests
Everyday messages, sounding like a person and StoryBench are our own benchmarks: blind, run in both orders, with a judge that is not one of the models being judged. EQ-Bench 4 is public and not ours. All the scores are on the model card.