There is no single best API for SillyTavern. It depends on what you play and how much has to stay private. For the widest choice of models behind one key, use OpenRouter; for privacy and no content filter, run a local backend such as KoboldCpp or Ollama on your own computer; for adult and dark roleplay from a hosted model, use an uncensored service such as Hemmingway Unlocked, which is ours and has a short list of things it never writes.
SillyTavern has no model of its own. It is an interface on your own machine, and the API you connect decides how well your characters write and what they will play. For the set-up steps, see our SillyTavern API guide.
The short list
- For the widest choice behind one key: OpenRouter. It puts models from many providers behind one OpenAI-compatible API, and SillyTavern has it as a built-in source.
- For privacy and no filter: a local backend, such as KoboldCpp or Ollama. The model runs on your own computer, and SillyTavern's docs describe local APIs as free to use with no content filter.
- For large open models without the hardware: Featherless. It hosts open-weight models behind one API key, and says it does not log what you send through its API.
- For co-writing stories: NovelAI. It is a subscription service built around storytelling with its own models.
- Ours, for adult and dark roleplay from a hosted model: Hemmingway Unlocked. It is uncensored and for adults only, with a short list of things it never writes, and it connects as a custom OpenAI-compatible address.
The options side by side
| Product | What it is | Good for | Watch out for |
|---|---|---|---|
| OpenRouter | One OpenAI-compatible API in front of models from many providers, paid per token | Trying many models with one key and one balance | Each provider behind it has its own policy on logging and retention |
| KoboldCpp | A single program that runs GGUF models on your own processor or graphics card | Privacy and no account: the chat stays on your machine | You need the hardware, and small models write weaker scenes |
| Ollama | A program that runs open models on your computer, with cloud models as well | The easiest local set-up, by SillyTavern's account | Only what runs locally stays on your machine |
| Featherless | Serverless hosting for open-weight models, many from Hugging Face | Open models too large to run at home | Which model you pick matters more than the service, so try a few |
| NovelAI | A subscription service for AI stories and anime images, with its own models | Co-writing prose, and it says prompts are encrypted | Its models continue text rather than follow instructions |
| Hemmingway Unlocked (ours) | The adults-only side of Hemmingway, with an uncensored model and an OpenAI-compatible API | Adult and dark roleplay without refusals or lectures | A short list of things it never writes, an API key needs a membership, and it writes text only |
| Hemmingway-1 API (ours) | Our regular model, hemmingway-27b | Scenes that are mostly conversation | Not the uncensored one, and larger story models write better long narration |
Three questions that decide it
- Must the chat stay on your computer? Then go local. SillyTavern's docs put the trade-off plainly: local APIs are free and unfiltered but take some setting up, and most local models are not as strong as cloud ones.
- Is it adult or very dark? By SillyTavern's account, cloud APIs filter content to varying degrees. On 27 September 2026 we sent the same adult roleplay request to ChatGPT, Claude and Gemini in their chat apps, and each refused; we did not test their APIs. The test is here. For adult scenes you want a local model or a hosted uncensored one.
- How do you want to pay? Per token, a flat subscription, a membership, or your own hardware. Every turn sends the character card and as much of the chat as fits, so per-token costs grow as a roleplay gets longer.
Hosted APIs
OpenRouter calls itself a unified API for the major models, with routing and fallbacks across providers, and you switch models without changing anything else. You connect it as a Chat Completion source, with OAuth or a pasted key. It is pay as you go, some models have a free variant with a daily limit, and you can leave out providers that train on your data.
Featherless gives you one API key and a catalogue of open-weight models. Its own SillyTavern guide connects it either as a Text Completion source or through the custom OpenAI-compatible source. It offers a flat-rate chat plan as well as per-token billing.
NovelAI says its story tool writes alongside you rather than for you and that prompts are encrypted. SillyTavern connects to it with a persistent API token from your NovelAI account. SillyTavern's docs add two cautions: the models continue the text they are given rather than follow instructions, so you switch off Instruct Mode, and replies are short.
Local backends
KoboldCpp is a single program with nothing to install. It runs GGUF models on your processor, your graphics card or both, and serves several kinds of API, an OpenAI-compatible one among them.
Ollama runs open models on your computer and says nothing you run locally leaves your machine. It offers cloud models too, which are not local.
The cost of going local is hardware. A model small enough for an ordinary laptop writes weaker scenes than a large hosted one. Open-source models for roleplay covers sizing a model to your memory.
Where Hemmingway fits
Hemmingway Unlocked is the adults-only side of Hemmingway: the same app with an uncensored model, at wise.hemmingway.io. It asks first whether you are 18 or over, then shows its rules. It writes sex between adults, explicit and in full, and dark, violent stories, without lectures or refusals on grounds of taste. A few things it never writes: nothing sexual involving anyone under 18, a real person, animals, family or anyone who does not want it, and never torture, help hurting or harassing a real person, weapons, making drugs, harmful software or help with self-harm. The list is in our terms, and every answer has a Report button.
Its API is OpenAI-compatible at https://wise.hemmingway.io/v1, with a key you create in Unlocked under the account menu, API keys. A key needs a membership; the free trial is for the app. Each message passes through our servers to be answered and is not kept, and chats are never trained on. It writes text only, and its weights are not out yet; we plan to open-source it soon. More on the Unlocked page.
Hemmingway-1, our regular model, is not the uncensored one. It was built for talking to people and for writing, not as a roleplay model. Its API is at https://hemmingway.io/v1 with the model hemmingway-27b. It placed third on the public EQ-Bench 4, a test of emotional intelligence, so it suits scenes that turn on conversation; larger story models scored higher on our StoryBench. Its thinking counts toward the reply limit, so do not set that limit low.
If nothing may leave your computer, run a local model instead: Unlocked's model cannot be downloaded yet.
Open Hemmingway UnlockedTry Hemmingway
Connecting any of them
For an OpenAI-compatible address, SillyTavern's docs give a short path: in API Connections choose Chat Completion, select Custom (OpenAI-compatible), enter the address and key, pick the model and press Test Message. If it will not connect, add /v1 to the address, never /chat/completions. The SillyTavern API guide has the rest.
How we chose
We picked one option per need rather than ranking them. We read each product's own site, and SillyTavern's docs for how each connects. We have not tested OpenRouter, KoboldCpp, Ollama, Featherless or NovelAI in SillyTavern, so we do not score them. Nor have we published a tested set-up for our own two APIs, so send a test message first. Hemmingway-1's scores are on the model card.
Common questions
What is the best API for SillyTavern?
It depends on what you play. OpenRouter gives the widest choice of models behind one key, a local backend such as KoboldCpp or Ollama keeps everything on your computer with no content filter, Featherless hosts open models without your own hardware, and NovelAI suits co-written stories. For adult roleplay from a hosted model, Hemmingway Unlocked, which is ours, connects as a custom OpenAI-compatible address.
Is there a free API for SillyTavern?
A local backend costs nothing beyond your own hardware. OpenRouter has some models with a free variant and a daily limit. Hemmingway Unlocked is free to try in its app, but an API key needs a membership.
Which SillyTavern API has no filter?
Local backends have no content filter, by SillyTavern's account, though a model can still be trained to refuse, and cloud APIs filter to varying degrees. Hemmingway Unlocked writes adult and dark fiction without refusing on grounds of taste, and still has a short list of things it never writes.
Is OpenRouter good for SillyTavern?
Yes, if you want to try many models without an account for each. SillyTavern has it as a built-in Chat Completion source. Most models are paid per token, and each provider behind it keeps its own logging policy.
Is a local model or an API better for SillyTavern?
A local model is private, unfiltered and costs nothing beyond your hardware, but most local models are weaker than hosted ones. An API gives you a stronger model with nothing to install, and your messages go to someone else's server. Choose local if privacy comes first and you have a strong graphics card.