You can usually tell. A message opens with "Certainly!", offers three versions when you asked for one, uses a word nobody says out loud, and closes by asking if you would like any changes. None of that is wrong, exactly. It just is not what a person does.
Hemmingway is an AI built to not do those things. This page is about what "sounds human" actually means, and how we measure it rather than just claim it.
What gives AI writing away
Most of it is not the words. It is the shape.
- The wrapper. You ask for a text to your landlord and get a preamble, the text, then a paragraph explaining the choices. A person would just send the text.
- Too many options. "Here are three ways you could say this." You wanted one.
- The register is off. A two-line reply to a friend comes back sounding like a cover letter.
- The tidy ending. A summary sentence, a question, an offer to help further.
- Made-up detail. A name, a date or a reason that was never in what you gave it, added to make the text feel complete.
A model can avoid every one of those and still read as a machine. But a model that does all of them never reads as a person.
How Hemmingway-1 was built for it
Hemmingway-1 is our own model: 27 billion parameters, built on Qwen3.8-27B, and trained on the writing people actually do every day. Messages, emails, the awkward note to a colleague, the thing you have been putting off. It was trained to give you the message and nothing else, in the register the situation calls for.
It is small next to the frontier models it is compared with below. That is on purpose: a model that only has to write well can be much smaller than one that has to do everything.
How we test "sounds human"
"Sounds human" is easy to say and hard to check. We check it three ways.
Human-Likeness
We took eighty real requests and had each model answer them. Every answer went head to head against another model's answer to the same request, and a judge model was asked one question: which of these two did a person write? The pairs were shuffled and run in both orders, so position could not sway the result, and the judge was never one of the models being judged.
| Model | Human-Likeness |
|---|---|
| Hemmingway-1 | 1032 |
| Fable 5.1 | 1006 |
| GLM-5.3 | 996 |
| Fable 5 | 992 |
| Kimi K3 | 985 |
| GPT-6 Astra | 964 |
| Qwen3.8 27B (base) | 952 |
Hemmingway-1 finished twenty-six points clear of the next model. This is our own benchmark, and we say so on the model card: we built it and ran it. The method is there for anyone to question.
EQ-Bench 4
EQ-Bench 4 is not ours. It is a public emotional-intelligence benchmark, and we ran it with its own harness. Hemmingway-1 placed third, behind Fable 5 and Kimi K3 and ahead of GPT-5.5, Opus 4.7 and Opus 4.8. Reading the room is a large part of sounding like a person. More on EQ-Bench 4 here.
Spot the AI
The last test is you. Spot the AI deals you two texts written for the same brief, one by a real person and one by Hemmingway, and asks you to pick the person. It is harder than it sounds.
Where it still falls short
It is not perfect, and a page like this should say where.
- It is English-first.
- It can be wrong and sound completely sure about it. Do not use it to decide anything medical, legal or financial.
- For long fiction, the big story models are better. In our StoryBench it sits level with Kimi K3, behind Fable 5 Max and GLM-5.3.
Trying it
Hemmingway runs in the browser, and as apps for Mac, Windows and Android. The apps can also read your mail and chats and draft the replies for you, in your voice. There is a free trial and no card is needed.