EQ-Bench is a public benchmark of emotional intelligence in language models. It puts a model into conversations that need it to read people: what someone actually needs, what they are not saying, how to answer without making things worse. A judge model then compares its replies with other models', and the results are ranked on an Elo scale.
It is not our benchmark, which is why we care about it. Our other tests are ones we built; this one we only ran.
Hemmingway-1's result
We ran EQ-Bench 4 on Hemmingway-1 with its official harness, against the leaderboard's own set of models.
| Model | EQ-Bench 4 |
|---|---|
| Fable 5 | 1341 |
| Kimi K3 | 1332 |
| Hemmingway-1 | 1330 |
| GPT-5.5 | 1316 |
| Opus 4.7 | 1312 |
| Opus 4.8 | 1285 |
That puts Hemmingway-1 third of the 27 models in the comparison, with a range of 1318 to 1340, inside twelve points of the top. It is a 27B model; every model around it is far larger.
Two things to know about the run
- We ran it ourselves. It was our run with the official harness and the leaderboard's models, not an entry the benchmark's maintainers ran for us.
- Thinking was on. Hemmingway-1 thinks before it answers by default, and it did here. The leaderboard turns reasoning off for most models it runs, so the comparison is close but not exact.
Why it matters for writing
A message that gets the words right and the person wrong still fails. The hardest things people ask a writing assistant for are exactly the ones EQ-Bench tests: turning someone down kindly, apologising without grovelling, answering a message that is angry about something else. This is also where Hemmingway-1 did best in our own tests. See those results.
All of Hemmingway-1's scores are on the model card.