The best AI for roleplay is one that stays in character, reads what the other person in the scene is feeling, leaves your character's words and actions to you, and holds the story together over a long conversation. For long, story-heavy roleplay, the large story models are the strongest we have tested: Fable 5 Max and GLM-5.3 score highest on our StoryBench. For roleplay that is mostly conversation, where the character has to read and answer a person well, emotional intelligence matters more, and there Fable 5, Kimi K3 and our own model, Hemmingway-1, sit at the top of the public EQ-Bench 4.
We build Hemmingway-1, so we should say plainly where it stands. It was built for talking to people and for writing, not as a roleplay model. It scores 1330 on EQ-Bench 4, third of the 27 models in the comparison. But in our own tests it loses on hostile storytelling and long story turns, and our README says outright that story models are better at those.
What makes a good roleplay model
Roleplay asks more of a model than most writing does. These are the things that decide it.
- Staying in character. The character should keep the same voice, the same history and the same wants from one reply to the next, and not drift into sounding like a helpful assistant.
- Reading the scene. A good partner notices what your character is feeling and what they are not saying, and answers that, not just the literal words.
- Not narrating over you. A common fault is the model writing your character's lines, deciding what you do next, or wrapping up a scene you wanted to keep open. A good model plays its part and hands the turn back.
- Keeping the story together. Over a long session the model has to remember who said what, where everyone is and what has already happened.
- Context length. This is how much text the model can hold at once. A longer context means more of the conversation can stay in view. How much of it a given app actually keeps is up to the app.
- Letting you steer. You should be able to set the tone, the pace and the length of replies, and have the model keep to them.
No single score measures all of that. Two kinds of benchmark come closest: one for reading people, and one for writing stories.
Reading people: EQ-Bench 4
EQ-Bench 4 is a public benchmark of emotional intelligence in conversation. It puts a model into conversations that need it to read people, and a judge compares its replies with other models' on an Elo scale. It is not our benchmark. We ran it ourselves with its official harness.
| Model | EQ-Bench 4 |
|---|---|
| Fable 5 | 1341 |
| Kimi K3 | 1332 |
| Hemmingway-1 | 1330 |
| GPT-5.5 | 1316 |
| Opus 4.7 | 1312 |
| Opus 4.8 | 1285 |
Fable 5 is first. Hemmingway-1 is third, inside twelve points of the top, and it is a 27B model among much larger ones. This is the score most relevant to roleplay that turns on how characters treat each other: a confession, a quarrel, a character comforting another. More on the EQ-Bench 4 run.
Writing stories: StoryBench
StoryBench is our own creative-writing benchmark. Each model writes stories to the same prompts, and every story goes head to head with another model's in a blind judgement, run in both orders, with a judge that is not one of the models being judged.
| Model | StoryBench |
|---|---|
| Fable 5 Max | 1277 |
| GLM-5.3 | 1254 |
| Hemmingway-1 | 1197 |
| Kimi K3 | 1197 |
| Qwen3.8 Max | 1081 |
| DeepSeek V4 Pro | 1054 |
Fable 5 Max and GLM-5.3 write better stories than Hemmingway-1 in our tests. Hemmingway-1 sits level with Kimi K3 and ahead of Qwen3.8 Max and DeepSeek V4 Pro. StoryBench tests stories written to a prompt, not a back-and-forth scene, so it is a guide to storytelling skill rather than a roleplay score. More on AI for creative writing.
Where Hemmingway-1 falls short
We break our own everyday-writing test down by the kind of request. It wins on money and admin, work, the hard asks people keep rewriting, and talking someone round. It loses on two that matter for roleplay: hostile storytelling and long story turns. In plain terms, when a scene needs a dark or antagonistic story told at length, or a long stretch of story in one reply, the story models do better. That is fair, and it is why we do not call Hemmingway-1 the best roleplay model.
Where it is more at home is the conversational side: a character who has to answer a person well, say something difficult kindly, or read a mood. That is the same skill it was built for in real messages, and the one EQ-Bench 4 tests.
Which to pick
- Long adventures, epic plots, lots of narration per turn: a large story model. Fable 5 Max and GLM-5.3 did best on our StoryBench.
- Character-driven scenes that turn on feelings: a model high on EQ-Bench 4. Fable 5, Kimi K3 and Hemmingway-1 are the top three in our comparison.
- An open model you can run yourself: Hemmingway-1 has open weights under Apache-2.0. Check each other model's own page for whether its weights are published. More on open models for roleplay.
- A ready-made character app: see Character.AI alternatives. We have not tested those apps.
How to get better roleplay from any model
Much of the difference comes from how you set the scene. These steps work with most chat models, including ours.
- Write a short character sheet before you start: name, age, how they speak, what they want, what they would never do.
- Say whose part the model plays and whose you play. For example: "You play Mara, the innkeeper. I play the traveller. Never write the traveller's words or actions."
- Set the length. "Keep each reply under 150 words and end on something the traveller can answer" stops the model running away with the scene.
- Set the tone in plain words: quiet, tense, funny, slow.
- In a long session, paste a short summary of what has happened so far every so often, so the important facts stay in view.
- If the model drifts out of character, say so directly and restate one line from the sheet.
A worked example of steps 2 and 3 together, as the first message:
You play Mara, who runs a roadside inn and trusts no one after dark. I play a traveller who arrives soaked at midnight. Write only Mara's words and actions, in the first person, under 120 words a reply, and stop when it is my turn.
Try HemmingwayDownload the app
About these numbers
StoryBench, and the category breakdown where Hemmingway-1 loses on hostile storytelling and long story turns, are our own tests. EQ-Bench 4 is public. All the scores are on the model card. We did not test Character.AI, companion apps, or models not listed above, so we cannot rank them.
Common questions
What is the best AI model for roleplay?
It depends on the kind of roleplay. For long, story-heavy scenes, Fable 5 Max and GLM-5.3 scored highest on our StoryBench. For character-driven conversation, Fable 5, Kimi K3 and Hemmingway-1 are the top three on the public EQ-Bench 4 in our comparison.
Is Hemmingway good for roleplay?
It is good at the conversational side: reading a mood and answering a person well, which is why it scores 1330 on EQ-Bench 4. It is weaker at hostile storytelling and long story turns, where story models do better in our tests. It was built for talking to people and writing, not as a roleplay model.
What is the best open-source model for roleplay?
In our comparison, GLM-5.3 is the strongest story writer on StoryBench after Fable 5 Max, and Kimi K3 is just ahead of Hemmingway-1 on EQ-Bench 4. Hemmingway-1 has open weights under Apache-2.0 and sits third on EQ-Bench 4. How to run one yourself.
How do I stop an AI from writing my character's lines?
Say it in the first message: whose part the model plays, whose part you play, and that it must never write your character's words or actions. Setting a word limit and asking it to stop when it is your turn helps too. If it slips, repeat the rule.
Does context length matter for roleplay?
Yes. Context length is how much of the conversation a model can hold at once, so a longer one keeps more of the story in view. Hemmingway-1's context is 262,144 tokens, though how much a given app sends to the model is up to the app.