We test the biggest models on writing all the time, because we build a writing model of our own and need to know where it stands. So we have blind results for Claude and ChatGPT's current models side by side. Here they are.
The models: Fable 5.1 and Fable 5 for Claude, and GPT-6 Astra for ChatGPT. EQ-Bench 4 compares the models in its own snapshot, so there it is Fable 5 against GPT-5.5. All runs were in September 2026.
The results
| Test | Claude | ChatGPT |
|---|---|---|
| Everyday messages | Fable 5.1 1024 | GPT-6 Astra 976 |
| Sounds like a person | Fable 5.1 1006 | GPT-6 Astra 964 |
| Long stories | Fable 5.1 1350 | GPT-6 Astra 1312 |
| EQ-Bench 4 (public) | Fable 5 1341 | GPT-5.5 1316 |
Scores are on an Elo scale: the higher score wins more head-to-head matchups. Claude came out ahead on every one.
Everyday messages and email
Fable 5.1 beat GPT-6 Astra by 48 points on eighty real email and message requests. But they lost in different ways.
- GPT-6 Astra is terse. Its messages were short, about 59 words on average, and it almost never wrapped them in commentary. What you get is usable as it is. It just reads less like a person wrote it.
- Fable 5.1 writes better, then wraps it. Its messages were the best quality of any model on the board, but about two in three came inside commentary, options or notes you have to cut before sending.
Stories
Fable 5.1 is the strongest story writer we have tested. In direct matchups it beat GPT-6 Astra 86 times to 22, with 11 ties. GPT-6 Astra is still a very good story writer, second only to Fable 5.1 on our board.
Emotional intelligence
On the public EQ-Bench 4, Fable 5 placed first in the comparison and GPT-5.5 fourth. Both read people well; Claude reads them a little better.
So which should you use?
- Long fiction: Claude, Fable 5.1.
- Short messages you want to paste and send without editing: ChatGPT's GPT-6 Astra is the more direct of the two.
- Messages that read like you wrote them: neither is the best at this. In the same tests, our own model Hemmingway-1 scored above both on everyday messages (1026) and on sounding like a person (1032), and gives you just the message. How it compares with Claude, and with ChatGPT.
About the tests
Everyday messages, sounding like a person and the story test are our own benchmarks. Each is blind: answers are shuffled, every matchup is run in both orders, and the judge is a model that is not one of those being judged. EQ-Bench 4 is public and not ours. We say which is which on our model card.