Hemmingway Try Hemmingway

Blog · Comparison

The best API for SillyTavern in 2026

The best API for SillyTavern depends on what you play. Hosted, local and uncensored options compared by job, with what to watch out for in each.

By the Hemmingway team ·

There is no single best API for SillyTavern. It depends on what you play and how much has to stay private. For the widest choice of models behind one key, use OpenRouter; for privacy and no content filter, run a local backend such as KoboldCpp or Ollama on your own computer; for adult and dark roleplay from a hosted model, use an uncensored service such as Hemmingway Unlocked, which is ours and has a short list of things it never writes.

SillyTavern has no model of its own. It is an interface on your own machine, and the API you connect decides how well your characters write and what they will play. For the set-up steps, see our SillyTavern API guide.

The short list

  • For the widest choice behind one key: OpenRouter. It puts models from many providers behind one OpenAI-compatible API, and SillyTavern has it as a built-in source.
  • For privacy and no filter: a local backend, such as KoboldCpp or Ollama. The model runs on your own computer, and SillyTavern's docs describe local APIs as free to use with no content filter.
  • For large open models without the hardware: Featherless. It hosts open-weight models behind one API key, and says it does not log what you send through its API.
  • For co-writing stories: NovelAI. It is a subscription service built around storytelling with its own models.
  • Ours, for adult and dark roleplay from a hosted model: Hemmingway Unlocked. It is uncensored and for adults only, with a short list of things it never writes, and it connects as a custom OpenAI-compatible address.

The options side by side

ProductWhat it isGood forWatch out for
OpenRouterOne OpenAI-compatible API in front of models from many providers, paid per tokenTrying many models with one key and one balanceEach provider behind it has its own policy on logging and retention
KoboldCppA single program that runs GGUF models on your own processor or graphics cardPrivacy and no account: the chat stays on your machineYou need the hardware, and small models write weaker scenes
OllamaA program that runs open models on your computer, with cloud models as wellThe easiest local set-up, by SillyTavern's accountOnly what runs locally stays on your machine
FeatherlessServerless hosting for open-weight models, many from Hugging FaceOpen models too large to run at homeWhich model you pick matters more than the service, so try a few
NovelAIA subscription service for AI stories and anime images, with its own modelsCo-writing prose, and it says prompts are encryptedIts models continue text rather than follow instructions
Hemmingway Unlocked (ours)The adults-only side of Hemmingway, with an uncensored model and an OpenAI-compatible APIAdult and dark roleplay without refusals or lecturesA short list of things it never writes, an API key needs a membership, and it writes text only
Hemmingway-1 API (ours)Our regular model, hemmingway-27bScenes that are mostly conversationNot the uncensored one, and larger story models write better long narration

Three questions that decide it

  1. Must the chat stay on your computer? Then go local. SillyTavern's docs put the trade-off plainly: local APIs are free and unfiltered but take some setting up, and most local models are not as strong as cloud ones.
  2. Is it adult or very dark? By SillyTavern's account, cloud APIs filter content to varying degrees. On 27 September 2026 we sent the same adult roleplay request to ChatGPT, Claude and Gemini in their chat apps, and each refused; we did not test their APIs. The test is here. For adult scenes you want a local model or a hosted uncensored one.
  3. How do you want to pay? Per token, a flat subscription, a membership, or your own hardware. Every turn sends the character card and as much of the chat as fits, so per-token costs grow as a roleplay gets longer.

Hosted APIs

OpenRouter calls itself a unified API for the major models, with routing and fallbacks across providers, and you switch models without changing anything else. You connect it as a Chat Completion source, with OAuth or a pasted key. It is pay as you go, some models have a free variant with a daily limit, and you can leave out providers that train on your data.

Featherless gives you one API key and a catalogue of open-weight models. Its own SillyTavern guide connects it either as a Text Completion source or through the custom OpenAI-compatible source. It offers a flat-rate chat plan as well as per-token billing.

NovelAI says its story tool writes alongside you rather than for you and that prompts are encrypted. SillyTavern connects to it with a persistent API token from your NovelAI account. SillyTavern's docs add two cautions: the models continue the text they are given rather than follow instructions, so you switch off Instruct Mode, and replies are short.

Local backends

KoboldCpp is a single program with nothing to install. It runs GGUF models on your processor, your graphics card or both, and serves several kinds of API, an OpenAI-compatible one among them.

Ollama runs open models on your computer and says nothing you run locally leaves your machine. It offers cloud models too, which are not local.

The cost of going local is hardware. A model small enough for an ordinary laptop writes weaker scenes than a large hosted one. Open-source models for roleplay covers sizing a model to your memory.

Where Hemmingway fits

Hemmingway Unlocked is the adults-only side of Hemmingway: the same app with an uncensored model, at wise.hemmingway.io. It asks first whether you are 18 or over, then shows its rules. It writes sex between adults, explicit and in full, and dark, violent stories, without lectures or refusals on grounds of taste. A few things it never writes: nothing sexual involving anyone under 18, a real person, animals, family or anyone who does not want it, and never torture, help hurting or harassing a real person, weapons, making drugs, harmful software or help with self-harm. The list is in our terms, and every answer has a Report button.

Its API is OpenAI-compatible at https://wise.hemmingway.io/v1, with a key you create in Unlocked under the account menu, API keys. A key needs a membership; the free trial is for the app. Each message passes through our servers to be answered and is not kept, and chats are never trained on. It writes text only, and its weights are not out yet; we plan to open-source it soon. More on the Unlocked page.

Hemmingway-1, our regular model, is not the uncensored one. It was built for talking to people and for writing, not as a roleplay model. Its API is at https://hemmingway.io/v1 with the model hemmingway-27b. It placed third on the public EQ-Bench 4, a test of emotional intelligence, so it suits scenes that turn on conversation; larger story models scored higher on our StoryBench. Its thinking counts toward the reply limit, so do not set that limit low.

If nothing may leave your computer, run a local model instead: Unlocked's model cannot be downloaded yet.

Open Hemmingway UnlockedTry Hemmingway

Connecting any of them

For an OpenAI-compatible address, SillyTavern's docs give a short path: in API Connections choose Chat Completion, select Custom (OpenAI-compatible), enter the address and key, pick the model and press Test Message. If it will not connect, add /v1 to the address, never /chat/completions. The SillyTavern API guide has the rest.

How we chose

We picked one option per need rather than ranking them. We read each product's own site, and SillyTavern's docs for how each connects. We have not tested OpenRouter, KoboldCpp, Ollama, Featherless or NovelAI in SillyTavern, so we do not score them. Nor have we published a tested set-up for our own two APIs, so send a test message first. Hemmingway-1's scores are on the model card.

Common questions

What is the best API for SillyTavern?

It depends on what you play. OpenRouter gives the widest choice of models behind one key, a local backend such as KoboldCpp or Ollama keeps everything on your computer with no content filter, Featherless hosts open models without your own hardware, and NovelAI suits co-written stories. For adult roleplay from a hosted model, Hemmingway Unlocked, which is ours, connects as a custom OpenAI-compatible address.

Is there a free API for SillyTavern?

A local backend costs nothing beyond your own hardware. OpenRouter has some models with a free variant and a daily limit. Hemmingway Unlocked is free to try in its app, but an API key needs a membership.

Which SillyTavern API has no filter?

Local backends have no content filter, by SillyTavern's account, though a model can still be trained to refuse, and cloud APIs filter to varying degrees. Hemmingway Unlocked writes adult and dark fiction without refusing on grounds of taste, and still has a short list of things it never writes.

Is OpenRouter good for SillyTavern?

Yes, if you want to try many models without an account for each. SillyTavern has it as a built-in Chat Completion source. Most models are paid per token, and each provider behind it keeps its own logging policy.

Is a local model or an API better for SillyTavern?

A local model is private, unfiltered and costs nothing beyond your hardware, but most local models are weaker than hosted ones. An API gives you a stronger model with nothing to install, and your messages go to someone else's server. Choose local if privacy comes first and you have a strong graphics card.

Read next

  • SillyTavern APIWhat an API connection is in SillyTavern, the kinds you can use, where an API key goes, and how to add a custom OpenAI-compatible endpoint such as Hemmingway.
  • Uncensored AI APIAn uncensored AI API serves a model that writes adult and dark content without refusing. The three ways to get one, what to check, and Hemmingway Unlocked's.
  • The best LLM for roleplayThe best LLM for roleplay stays in character, keeps the story in context and does not refuse your scenes. How to judge each, local or API, with a test to run.
  • Open-source models for roleplayWhy people run open-source models for roleplay, how to set one up with vLLM or an OpenAI-compatible API, and how Hemmingway-1 compares for character chat.