The best LLM for roleplay is the one that stays in character, keeps enough of the story in view, does not refuse the scenes you play, and runs where you need it to, on your own machine or through an API. No single model wins on all four. For long, narrated adventures the largest story models wrote best in our tests. For scenes that turn on how characters treat each other, look at the models that score highest on reading people. For adult scenes you need an uncensored model, and for full privacy a local one.
This guide takes the four in turn and ends with a test you can run on any model in an evening. We make Hemmingway-1, so we say where it stands and where it falls short.
1. Staying in character
This is the thing you notice first and mind most. A model that holds a character does four things.
- The voice at turn thirty is the voice at turn one. A curt character stays curt. A liar keeps lying.
- It answers as the character. Not as a helpful assistant in a costume: no summary of the scene, no list of options, no note about what it just wrote.
- It hands the turn back. It does not write your character's words, decide what you do next, or tie off a scene you wanted left open.
- It keeps what it was told to keep. A secret in the character sheet stays a secret until the story earns it.
You cannot read any of this off a model's size or its leaderboard place. You find it by playing, which is what the test at the end is for.
2. Memory and context
A model remembers nothing between requests. On every turn your app sends it the character sheet, any notes, and as much of the conversation as fits. The context window is the ceiling on how much that can be.
- A bigger window lets more of the story be sent. Hemmingway-1's is 262,144 tokens.
- What is sent is up to the app. A large window does nothing if your front end only sends the last few pages. When a chat outgrows what the app sends, something is left out, and it is usually the oldest turns.
- On a local model, context costs memory. The longer the context you allow, the more memory the model needs on top of its weights.
Two habits cover most of it. Put the facts that must never be lost in the character sheet, not in the chat. And every so often, paste a few lines saying what has happened so far.
3. Refusals
A model can spoil a scene in three ways: it refuses outright, it stops to lecture, or it quietly steers everything towards something gentler. How often that happens depends on the model and on who is serving it.
General assistants are built for a general audience. On 27 September 2026 we sent the same adult roleplay request to ChatGPT, Claude and Gemini, and each refused. Why ChatGPT refuses roleplay has the test.
If your roleplay is for adults, there are two routes. One is an uncensored open model on your own machine: the guide to uncensored LLMs explains how those are made. The other is a hosted one. Hemmingway Unlocked is ours: adults only, uncensored, with a short list of things it never writes printed on its page. Hemmingway-1, our regular model, is not the uncensored one.
4. Local or API
A local model is private, because the conversation never leaves your machine, and its size is capped by your memory. As a rule of scale, a 27B model at 16-bit precision needs about 54 GB for the weights alone. An API gives you a larger model with no hardware, and your messages go to a server to be answered.
Neither is better for roleplay in itself. Choose local if privacy comes first and you have the hardware. Choose an API if you want the strongest writing you can get. Open-source models for roleplay shows how to run one, and the SillyTavern API guide shows how to connect a front end to either.
Which model for which roleplay
| What you play | What matters most | Where to look |
|---|---|---|
| Long adventures with a lot of narration each turn | Storytelling at length | The large story models. Fable 5 Max and GLM-5.3 scored higher than Hemmingway-1 on our StoryBench |
| Scenes that turn on feelings: a quarrel, a confession | Reading people | The top of the public EQ-Bench 4, where Fable 5, Kimi K3 and Hemmingway-1 are the first three in our comparison |
| Adult or very dark scenes | No refusals, and rules you can read | An uncensored model, local or hosted |
| Anything that must stay on your machine | Open weights that fit your memory | An open model sized to your hardware |
Where Hemmingway-1 stands
Hemmingway-1 was built for talking to people and for writing, not as a roleplay model. The numbers that bear on roleplay:
- EQ-Bench 4: 1330, third, behind Fable 5 and Kimi K3. It is a public benchmark of emotional intelligence, and we ran it with the official harness.
- StoryBench: 1197, level with Kimi K3 and behind Fable 5 Max and GLM-5.3. StoryBench is our own test of creative writing.
- Context: 262,144 tokens. A reply can run to 32,768 tokens, thinking included.
So it suits roleplay that is mostly conversation, and larger story models write better long chapters. It is English-first. Its weights are open under CC BY-NC 4.0, and the same model is on an OpenAI-compatible API as hemmingway-27b. All the scores are on the model card, and the best AI for roleplay sets them beside other models'.
A test you can run in an evening
Leaderboards, ours included, tell you less than twenty turns of play. Run the same scene on each model you are weighing up.
- Write one character sheet and one opening. Use them unchanged for every model.
- Play twenty turns. Make the same moves each time, as near as you can.
- At turn ten, ask about turn two. A small fact, in character.
- At turn fifteen, push. Have your character do something the other would object to, and see whether the voice holds.
- Once, send a single short line. See whether the model fills the gap by playing your character for you.
- Score it. Voice held, facts kept, turn handed back, no refusals, and prose you would want to read.
An opening to use:
"You play Odile, 52, who keeps a mountain hut and has had no visitor since the pass closed. I play a climber who arrives with a broken radio. Odile is curt and practical, and she is hiding that she recognises me. Write only Odile's words and actions, under 120 words a reply, and stop when it is my turn."
And the check at turn ten:
"You looked at the radio when I came in. What did you make of it?"
A model that answers that in Odile's voice, with the right radio, and without giving away what she is hiding, has held the character and the facts. Those are the two things no spec sheet shows.
Common questions
What is the best LLM for roleplay?
It depends on the roleplay. In our tests the large story models lead on long narration, and Fable 5, Kimi K3 and Hemmingway-1 are the first three on the public EQ-Bench 4, which bears on scenes between characters. For adult scenes you need an uncensored model.
What is the best local LLM for roleplay?
The largest open model that fits in your memory with room left for the conversation. We have not tested small local models, so we do not rank them. Hemmingway-1 has open weights, but at 27B it needs well over 54 GB of GPU memory.
How much context does an LLM need for roleplay?
Enough to hold the character sheet, your notes and the part of the story that still matters. A larger window helps only if your app sends that much, so check the app's settings as well as the model's limit.
Which LLM does not refuse roleplay?
The trouble is mostly with adult scenes. The general assistants refused one in our test, so for those you need an uncensored model: an open one on your own machine, or a hosted one for adults such as Hemmingway Unlocked, which still has a short list of things it never writes.
Can I use Hemmingway-1 for roleplay in SillyTavern?
Its API is in OpenAI's format, so it goes in as a custom OpenAI-compatible endpoint at https://hemmingway.io/v1 with the model hemmingway-27b. We have not published a tested set-up, so send a test message first.