Hemmingway Try Hemmingway

Blog · Guide

The best AI for creative writing in 2026

Which AI models write the best stories, poems and letters? Our blind StoryBench and EQ-Bench 4 results, including where our own model loses.

By the Hemmingway team ·

There is no single best model for creative writing, because "creative writing" covers a four-line poem, a wedding speech and a ninety-thousand-word novel. So this page gives you our test results and says plainly which model is strongest at what, including where our own model, Hemmingway-1, is not the one to pick.

Long stories: StoryBench

StoryBench is our own story test. Each model writes stories to the same prompts, and every story goes head to head with another model's in a blind judgement, run in both orders. Scores are on an Elo scale: higher wins more often.

ModelStoryBench
Fable 5 Max1277
GLM-5.31254
Hemmingway-11197
Kimi K31197
Qwen3.8 Max1081
DeepSeek V4 Pro1054
Qwen3.8 27B (base)693

For long fiction, the big story models win. Fable 5 Max and GLM-5.3 write better long stories than Hemmingway-1 in our tests, and so do the newest frontier models. Hemmingway-1 sits level with Kimi K3 and well ahead of Qwen3.8 Max and DeepSeek V4 Pro, and 504 points above the model it was built from. For a 27B model that is a lot, but it is not first.

Reading people: EQ-Bench 4

A lot of creative writing is getting a character's feelings right. EQ-Bench 4 is a public emotional-intelligence benchmark, and we ran it with its official harness.

ModelEQ-Bench 4
Fable 51341
Kimi K31332
Hemmingway-11330
GPT-5.51316
Opus 4.71312
Opus 4.81285

Hemmingway-1 is third, inside twelve points of the top.

Short pieces that have to sound like a person

This is where Hemmingway-1 is strongest: letters, speeches, messages, short stories and poems that need to read as if someone sat down and wrote them. In our Human-Likeness test, where a judge picks which of two texts a person wrote, it came first, twenty-six points ahead of the next model.

Which to pick

  • A novel, or long chapters: one of the large story models. Fable 5 Max did best in our tests.
  • Letters, speeches, cards, short stories, poems: Hemmingway, which is built to sound like a person and keeps to the length you asked for.
  • Something you will run yourself: Hemmingway-1 is open weights under Apache-2.0. More on open models for writing.

About these numbers

StoryBench and Human-Likeness are our own benchmarks, and we say so on the model card. Every matchup is blind, run in both orders, and judged by a model that is not one of the ones being judged. EQ-Bench 4 is not ours.

A sample

Here is a story Hemmingway-1 wrote, unedited, for the prompt "Write a story of about 150 words about a lighthouse keeper who gets a letter meant for someone else."

The supply boat came up the rock in November rain, and the boy who brought the stores had one more thing: a letter, addressed to Mr. Brannan, keeper of Vane Rock Light.

Callum took it without opening it, as he always did. After the lamp was lit and the log signed, he broke the seal. It was from a woman in Oban, Brannan's sister. She said she was leaving the mainland in spring. She said the things she had not said in the autumn were said now, all of them, and that she did not expect an answer.

Callum had kept letters like this for thirty years, for other men, other lights. None had ever been sent on. He read hers twice, then folded it along its old creases.

At four bells he climbed the stairs and trimmed the wick. The beam went out over the water, and the letter stayed in the drawer.

Try HemmingwayDownload the app

Read next

  • The best open-source LLM for writingHemmingway-1 is a 27B open-weights model under Apache-2.0, built for writing that sounds like a person. How it scores, what it needs, and how to run it.
  • EQ-Bench 4: how Hemmingway-1 scoresEQ-Bench 4 tests emotional intelligence in conversation. We ran it on Hemmingway-1 with the official harness: 1330, third of 27, ahead of GPT-5.5 and Opus.
  • Claude vs ChatGPT for writingWhich writes better, Claude or ChatGPT? Our blind tests of Fable 5.1 and GPT-6 Astra on everyday messages, sounding human, stories and emotional intelligence.
  • An AI that sounds humanMost AI writing gives itself away in a line. Here is what makes text read as a person's, and how Hemmingway-1 was built and tested to sound like one.