Hemmingway Try Hemmingway

Blog · Guide

EQ-Bench 4: how Hemmingway-1 scores

EQ-Bench 4 tests emotional intelligence in conversation. We ran it on Hemmingway-1 with the official harness: 1330, third of 27, ahead of GPT-5.5 and Opus.

By the Hemmingway team ·

EQ-Bench is a public benchmark of emotional intelligence in language models. It puts a model into conversations that need it to read people: what someone actually needs, what they are not saying, how to answer without making things worse. A judge model then compares its replies with other models', and the results are ranked on an Elo scale.

It is not our benchmark, which is why we care about it. Our other tests are ones we built; this one we only ran.

Hemmingway-1's result

We ran EQ-Bench 4 on Hemmingway-1 with its official harness, against the leaderboard's own set of models.

ModelEQ-Bench 4
Fable 51341
Kimi K31332
Hemmingway-11330
GPT-5.51316
Opus 4.71312
Opus 4.81285

That puts Hemmingway-1 third of the 27 models in the comparison, with a range of 1318 to 1340, inside twelve points of the top. It is a 27B model; every model around it is far larger.

Two things to know about the run

  • We ran it ourselves. It was our run with the official harness and the leaderboard's models, not an entry the benchmark's maintainers ran for us.
  • Thinking was on. Hemmingway-1 thinks before it answers by default, and it did here. The leaderboard turns reasoning off for most models it runs, so the comparison is close but not exact.

Why it matters for writing

A message that gets the words right and the person wrong still fails. The hardest things people ask a writing assistant for are exactly the ones EQ-Bench tests: turning someone down kindly, apologising without grovelling, answering a message that is angry about something else. This is also where Hemmingway-1 did best in our own tests. See those results.

All of Hemmingway-1's scores are on the model card.

Try HemmingwayDownload the app

Read next

  • An AI that sounds humanMost AI writing gives itself away in a line. Here is what makes text read as a person's, and how Hemmingway-1 was built and tested to sound like one.
  • The best AI for creative writing in 2026Which AI models write the best stories, poems and letters? Our blind StoryBench and EQ-Bench 4 results, including where our own model loses.
  • The best open-source LLM for writingHemmingway-1 is a 27B open-weights model under Apache-2.0, built for writing that sounds like a person. How it scores, what it needs, and how to run it.
  • Claude vs ChatGPT for writingWhich writes better, Claude or ChatGPT? Our blind tests of Fable 5.1 and GPT-6 Astra on everyday messages, sounding human, stories and emotional intelligence.