There is no single best AI for writing, because writing is several different jobs. In our tests, the best for everyday messages and emails, and for writing that sounds like a person, is our own model, Hemmingway-1. The best for long fiction are the large story models, led by Fable 5 Max and GLM-5.3. For checking and polishing your own draft, a grammar checker such as Grammarly does a different job from any of these. For marketing copy, tools built for marketing teams, such as Jasper, are the obvious place to look.
We make Hemmingway, so we test the big models against each other all the time. This page gives the results by job, says which tools we did not test, and says where our model loses. Three of the tests are our own benchmarks, and we mark them as ours every time.
The short answer
| Job | Try first | Tested by us |
|---|---|---|
| Everyday messages and emails | Hemmingway | Yes |
| Sounding like a person | Hemmingway | Yes |
| Long fiction and novels | A large story model, such as Fable 5 Max or GLM-5.3 | Yes |
| Short creative pieces | Hemmingway | Yes |
| Editing and grammar | A checker, such as Grammarly | No |
| Marketing copy | A marketing tool, such as Jasper | No |
Everyday messages and emails
Most writing is not a novel. It is the reply to a client, the text to a landlord, the email you have been putting off. CommunicationBench is our own test of exactly that: eighty real requests, every answer judged blind against another model's answer, in both orders.
| Model | CommunicationBench (ours) |
|---|---|
| Hemmingway-1 | 1026 |
| Fable 5.1 | 1024 |
| Fable 5 | 1009 |
| GLM-5.3 | 1007 |
| Kimi K3 | 996 |
| GPT-6 Astra | 976 |
| Qwen3.8 27B (base) | 954 |
Hemmingway-1 came first, level with Fable 5.1 and fifty points ahead of GPT-6 Astra, the model behind ChatGPT in our tests. It is a 27B model; most of the others are far larger. It also tends to give you the message itself rather than a set of options with notes attached. Fable 5, GLM-5.3 and Kimi K3 wrapped their answer in commentary in more than nine replies out of ten.
The gap is widest on hard asks: saying no, asking for money back, telling someone something they will not like. Judged on which reply sounded like a person wrote it, GPT-6 Astra's won 9% of the time and Hemmingway-1's 72%.
More on this job: the best AI for emails and messages, AI email replies, AI text message replies, how to reply to an angry email and how to decline politely.
Sounding like a person
Human-Likeness is our own test built on the same matchups, with one question for the judge: which of these two did a person write?
| Model | Human-Likeness (ours) |
|---|---|
| Hemmingway-1 | 1032 |
| Fable 5.1 | 1006 |
| GLM-5.3 | 996 |
| Fable 5 | 992 |
| Kimi K3 | 985 |
| GPT-6 Astra | 964 |
Hemmingway-1 finished twenty-six points clear of the next model. If the problem with AI writing for you is that it reads like AI, this is the number that matters. You can test it yourself at Spot the AI.
More on this job: AI that sounds human, how to make AI text sound human and how to make AI write like you.
Creative writing and long fiction
StoryBench is our own creative-writing test: stories written to the same prompts, judged blind in both orders.
| Model | StoryBench (ours) |
|---|---|
| Fable 5 Max | 1277 |
| GLM-5.3 | 1254 |
| Hemmingway-1 | 1197 |
| Kimi K3 | 1197 |
| Qwen3.8 Max | 1081 |
| DeepSeek V4 Pro | 1054 |
For long fiction, Hemmingway-1 is not the pick. Fable 5 Max and GLM-5.3 write better stories in our tests. Hemmingway-1 sits level with Kimi K3, well ahead of Qwen3.8 Max and DeepSeek V4 Pro. It is at home with short pieces that have to sound like a person: a letter, a speech, a card, a short story or a poem. We did not test tools built for novelists, such as Sudowrite.
Feelings on the page matter in fiction too. On the public EQ-Bench 4, which tests emotional intelligence in conversation and which we ran with its official harness, Fable 5 scored 1341, Kimi K3 1332 and Hemmingway-1 1330, ahead of GPT-5.5, Opus 4.7 and Opus 4.8.
More on this job: the best AI for creative writing, Hemmingway vs Sudowrite, AI story writer, AI poem writer, AI wedding speech and the best AI for roleplay.
Try HemmingwayDownload the app
Editing and grammar
If you like writing your own words and want them cleaner, you need a different kind of tool: one that starts from your draft rather than from what you want to say. We did not test these, so here is what each is for, not which is best.
- Grammarly checks spelling, grammar, clarity and tone as you type, with AI suggestions for rewriting.
- Hemingway Editor, with one m, highlights long and hard-to-read sentences so you can fix them yourself. It is a different product from a different company, not connected to us.
- QuillBot is best known for paraphrasing: it rewords a sentence or paragraph you give it.
Any writing assistant, ours included, will also proofread if you paste in a draft and ask. More on this job: Grammarly alternatives, Hemmingway vs Grammarly, Hemmingway vs Hemingway Editor and Hemmingway vs QuillBot.
Marketing copy
Marketing content for a company, with a brand voice that has to stay the same across a team, is its own job. Jasper is built for marketing teams, and Hemmingway is not. We have not tested Jasper. Hemmingway is built around your own voice, which changes with the person you write to, rather than a company's. Hemmingway vs Jasper goes into the difference.
General assistants
ChatGPT, Claude and Gemini are general assistants: they write, answer questions, help with code and much more. If you want one tool for everything, one of them is a sensible choice. In our tests, Claude's Fable models are strong at long fiction, and GPT-6 Astra's messages are short and usable as they are, but sound less like a person than Hemmingway-1's. We did not benchmark Gemini. More: ChatGPT alternatives for writing, Claude vs ChatGPT for writing, Hemmingway vs ChatGPT, Hemmingway vs Claude and Hemmingway vs Gemini.
If you want to run it yourself
Hemmingway-1 is open weights under Apache-2.0: 27B, built on Qwen3.8-27B, with a 262,144-token context. You can run it yourself or use it through an OpenAI-compatible API as hemmingway-27b. See the best open-source LLM for writing, how to run Hemmingway-1 locally and the best LLM API for writing.
About these numbers
CommunicationBench, Human-Likeness and StoryBench are our own benchmarks. Every matchup was blind, run in both orders, and judged by a model that was not one of those being judged. EQ-Bench 4 is public. All scores are on the model card. Hemmingway-1 is English-first, and it can be wrong and still sound certain, so do not use it for medical, legal or money decisions.
Common questions
What is the best AI for writing?
It depends on the job. In our tests, Hemmingway-1 was best at everyday messages and emails and at sounding like a person, while Fable 5 Max and GLM-5.3 were best at long fiction. For checking your own draft, a grammar checker does a different job. Our three writing tests are our own benchmarks.
What is the best AI for writing emails?
In our CommunicationBench, Hemmingway-1 scored 1026, level with Fable 5.1 and ahead of Fable 5, GLM-5.3, Kimi K3 and GPT-6 Astra. CommunicationBench is our own benchmark. More on emails and messages.
What is the best AI for writing a novel?
A large story model. On our StoryBench, Fable 5 Max scored 1277 and GLM-5.3 1254, ahead of Hemmingway-1 at 1197. We did not test novel-writing tools such as Sudowrite.
Which AI writing sounds the most human?
In our Human-Likeness test, where a judge picks which of two texts a person wrote, Hemmingway-1 came first with 1032, twenty-six points ahead of Fable 5.1. Human-Likeness is our own benchmark. You can try telling them apart yourself at Spot the AI.
Is there a free AI for writing?
Hemmingway has a free trial with no card needed, in the apps and at /app/. For what the plans include, see pricing.