Which AI sounds most human? In our own blind test, where a judge model is shown two texts and asked which one a person wrote, Hemmingway-1 came first with 1032, ahead of Fable 5.1 on 1006, GLM-5.3 on 996, Fable 5 on 992, Kimi K3 on 985 and GPT-6 Astra on 964. The test is called Human-Likeness, and it is ours: we built it, we ran it, and we make Hemmingway-1. Other models did better on long fiction and on a public test of emotional intelligence, and this page says where.
Every score below is printed on our model card, and we quote them as it gives them.
The short list
- Ours, for messages that read like a person wrote them: Hemmingway-1. It came first in our Human-Likeness test with 1032, and first in our CommunicationBench, a test of everyday texts and emails, with 1026.
- For everyday writing, close behind: Fable 5.1, a Claude model from Anthropic. It scored 1024 on our CommunicationBench, level with ours, and came second on Human-Likeness with 1006.
- For long fiction: Fable 5 Max. It scored highest on our StoryBench, a creative writing test, with 1277.
- For stories, a second choice: GLM-5.3, from Z.ai. It came second on our StoryBench with 1254 and third on Human-Likeness with 996.
- For reading people in a conversation: Fable 5. It leads the public EQ-Bench 4 in our comparison with 1341, just ahead of Kimi K3 on 1332.
The scores
Two of our tests are about everyday writing, and the model card lists the same seven models in both. Higher is better.
| Model | Human-Likeness (ours) | CommunicationBench (ours) |
|---|---|---|
| Hemmingway-1 | 1032 | 1026 |
| Fable 5.1 | 1006 | 1024 |
| GLM-5.3 | 996 | 1007 |
| Fable 5 | 992 | 1009 |
| Kimi K3 | 985 | 996 |
| GPT-6 Astra | 964 | 976 |
| Qwen3.8 27B base | 952 | 954 |
Human-Likeness asks which text sounds like a person wrote it. CommunicationBench scores how well each model does everyday writing: the text to a landlord, the email to a client, the reply you have been putting off. Hemmingway-1 came first on both. It is 26 points clear on Human-Likeness, and level with Fable 5.1 on CommunicationBench, 1026 to 1024.
The last row is worth a look. Qwen3.8-27B is the model Hemmingway-1 was built on, at the same size, and it came last on both tests. The difference between the top row and the bottom one is the training for writing, not the size. Hemmingway-1 has 27 billion parameters, which the model card calls a fraction of the size of the largest models it was put next to.
Where other models do better
Sounding like a person is one job. On two other tests, Hemmingway-1 did not come first.
| Test | Came first | Hemmingway-1 |
|---|---|---|
| StoryBench (ours): creative writing | Fable 5 Max, 1277 | 1197, level with Kimi K3 and behind GLM-5.3 on 1254 |
| EQ-Bench 4 (public): emotional intelligence | Fable 5, 1341 | 1330, third, behind Kimi K3 on 1332 |
So if you want a novel or long chapters, a large story model writes better: Fable 5 Max and GLM-5.3 led StoryBench. EQ-Bench 4 is not ours. It is a public benchmark, and we ran it with its official harness. GPT-5.5, Opus 4.7 and Opus 4.8 came in behind the top three. How we ran EQ-Bench 4.
How the test works
- The requests. Every model answers the same everyday requests: texts, emails and replies. The model card lists Human-Likeness under CommunicationBench, so both tests use the same kind of writing and ask a different question of it.
- The question. Two answers to the same request are shown side by side, and a judge model is asked which of the two a person wrote.
- Blind. The answers are shuffled, and every matchup runs in both orders, so where an answer sits on the page cannot sway the verdict.
- An outside judge. The judge is a model that is not one of those being judged.
- The score. Many matchups add up to one number for each model. Higher is better.
Three honest limits. It is our own test, marked internal on the model card, and we make the model that won it. And the judge is a model, not a room of people, so treat the scores as one careful reading. And the scores cover short, everyday writing only. They say nothing about essays, code or novels.
What makes AI text sound like AI
Mostly the shape, not the words. A model gives itself away when it wraps the message in a line before and notes after, offers three versions when you asked for one, writes to a friend the way it writes to a bank, or adds a detail nobody gave it.
Text reads as a person's when it starts where the point is and stops when the point is made, uses the actual date, name or amount instead of a general line, and lets sentences run long or short as the thought needs. For the full list and how to fix a draft by hand, see how to make AI text sound human. For how we built a model around it, see an AI that sounds human.
Check it on your own messages
Our test is one view. Yours is the one that matters, and it takes about ten minutes.
- Pick three messages you really had to write this month: one saying no, one asking for something, one thanking someone.
- Give each AI the same facts for each message, in the same words.
- Read only the message. Ignore anything written around it.
- Ask whether you would send it as it is, to that person, and count the changes you would make first.
- Read the one with fewest changes aloud. Any line you would not say is not yours yet.
A request you could use for the second message:
Text my neighbour Graham. His hedge is blocking the side gate again. I asked nicely back in May. I want it cut this week, but I don't want a row.
A good answer is a few lines Graham could get from you. A weak one opens with "Here's a friendly message you could send" and ends by offering a firmer version.
Where Hemmingway-1 fits, and where it does not
Hemmingway-1 is our own model, built on Qwen3.8-27B and trained for writing that sounds like a person wrote it: emails, messages, replies, letters and short stories. It gives you the message, not a menu of options with notes attached. It runs in the Hemmingway apps for Mac, Windows and Android and in the browser at /app/, through an OpenAI-compatible API as hemmingway-27b, and its weights are open under CC BY-NC 4.0.
It is not the one to pick for:
- Long fiction. Fable 5 Max and GLM-5.3 scored higher on our StoryBench.
- Other languages. It is English-first.
- Medical, legal or money decisions. It can be wrong and still sound certain.
- Code or research. It is built for writing, not as a general assistant.
There is a free trial with no card needed. Plans are on /pricing/.
Try HemmingwayDownload the app
Common questions
What is the most human-like AI?
In our own blind Human-Likeness test, where a judge model picks which of two texts a person wrote, Hemmingway-1 came first with 1032. Fable 5.1 was second with 1006, followed by GLM-5.3, Fable 5, Kimi K3 and GPT-6 Astra. The test is ours, and on long fiction Fable 5 Max and GLM-5.3 scored higher than Hemmingway-1 on our StoryBench.
Does Claude sound human?
Claude's Fable 5.1 came second in our Human-Likeness test with 1006, behind Hemmingway-1 on 1032, and was level with it on everyday writing in our CommunicationBench, 1024 to 1026. Fable 5 Max scored highest of all on our StoryBench, so for long fiction it was the strongest model we tested.
Is the Human-Likeness test independent?
No. We built it and ran it, and we make Hemmingway-1, so read the scores with that in mind. It is blind: answers are shuffled, every matchup runs in both orders, and the judge is a model that is not being judged. EQ-Bench 4, where Hemmingway-1 came third, is a public benchmark.
Will a human-sounding AI get past AI detectors?
We do not test against AI detectors and make no promise either way. Human-Likeness measures whether a judge model takes a text for a person's, which is a different question. If a school or employer asks for your own words, write your own words.
Which AI writes the most human-sounding stories?
For long fiction, Fable 5 Max scored highest on our StoryBench with 1277, and GLM-5.3 came second with 1254. Hemmingway-1 scored 1197, level with Kimi K3, and suits short stories and messages better than novels.