SillyTavern has no model of its own. It is a front end: it builds a prompt from your chat, your character and your settings, and sends it to a model through an API connection that you set up. That connection can be a model running on your own computer, a hosted service you have a key for, or any endpoint that speaks OpenAI's format. You choose it in the API Connections panel, which is the second button in the top bar.
This guide covers what the choices mean, where an API key goes, and how to add a custom OpenAI-compatible endpoint. We make Hemmingway, so the last part shows its two addresses. The steps are the same for any compatible API. What we say about SillyTavern's menus comes from its own docs at docs.sillytavern.app.
What an API connection is
Two settings decide how SillyTavern talks to a model.
- The API type. SillyTavern's docs describe two main ones. Chat Completion sends the prompt as a list of messages between the user, the assistant and the system. Text Completion turns the whole prompt into one long string, and the model continues it. With Text Completion you can apply an instruct template, which helps the model respond to instructions.
- The source or backend. This is who answers: a program on your machine, or a service on the internet.
An OpenAI-compatible address is always a Chat Completion connection.
The kinds of API
| Kind | Examples in SillyTavern's docs | Good for | Watch out for |
|---|---|---|---|
| Local backends | KoboldCpp, llama.cpp, Ollama, Oobabooga's TextGeneration WebUI, TabbyAPI | Privacy, and no content filter | You download the models and need the hardware to run them |
| Hosted APIs | OpenAI, Claude, Google AI Studio, Mistral, OpenRouter, DeepSeek, NovelAI | Stronger models with nothing to install | Content filtering of varying degrees, and most are paid |
| Custom OpenAI-compatible | LM Studio, LiteLLM, LocalAI, or any service with a compatible address | Anything not on the built-in list | SillyTavern does not promise every endpoint works |
The docs sum up the trade-off in much the same words: local models are free to use and unfiltered, and most are not as strong as the hosted ones. Hosted models need nothing from your computer and are stronger, and all of them filter content to some degree.
API keys
A key is how a hosted service knows a request is yours and which account to charge.
- Where it comes from. Each service issues its own, on its own site. SillyTavern does not give you one.
- Where it goes. In the API Connections panel, in the key field for the source you picked.
- Local backends. A program on your own machine usually needs no key. Ollama's docs say clients must send a key value, but Ollama ignores it.
- Keep it private. A key is as good as a password for spending your credit. Do not paste it into a shared preset, a screenshot or a forum post.
If you switch between models often, SillyTavern's Connection Profiles save the API type, the model, the server address and the key together, so changing model is one choice in a list. Set the connection up, click Create, and name it.
How to add a custom OpenAI-compatible endpoint
These are the steps from SillyTavern's docs.
- Open API Connections, the second button in the top bar.
- Set the API type to Chat Completion.
- For Chat Completion Source, select Custom (OpenAI-compatible).
- Enter the endpoint address, and the API key if the service needs one.
- Choose the model. If the endpoint lists its models, you pick from a dropdown. If it does not, type the model ID into the text field.
- Click Test Message. It sends a short prompt to check the connection.
Two details save most of the trouble. If it will not connect, add /v1 to the end of the address. And do not add /chat/completions to it.
There is also an option called Bypass API status check, which stops SillyTavern warning you that the endpoint is not working. And under Prompt Post-Processing you can have SillyTavern reshape the messages before sending: merge consecutive messages from the same role, allow only one system message, require a user message first, or fold everything into a single user message. Leave it at None unless an endpoint complains about the order of messages.
SillyTavern's docs are clear that they give no support for custom endpoints and no promise that every one works. Test before a long session.
Connecting a local backend
For KoboldCpp, SillyTavern's guide to self-hosted models gives these settings: API set to Text Completion, API Type set to KoboldCpp, the server address http://127.0.0.1:5001/, then Connect.
LM Studio and Ollama can also be reached as custom OpenAI-compatible endpoints. In LM Studio you start the server in the Developer tab, and its docs use http://localhost:1234/v1 as the address. Ollama's docs give http://localhost:11434/v1/.
For how much memory a model needs, and how to serve Hemmingway-1's open weights yourself, see open-source models for roleplay.
Using Hemmingway in SillyTavern
Hemmingway's API takes OpenAI's Chat Completions body, so it goes in as a custom endpoint.
- Address:
https://hemmingway.io/v1 - Key: one you make on the API platform, where the prices are too. Keys start with
hemmingway_live_. - Model:
hemmingway-27b
Three things to know. The model thinks before it answers, and the thinking arrives in its own field, reasoning_content, apart from the reply. The reply limit counts the thinking too, so do not set it very low. And the model was built for talking to people and for writing, not as a roleplay model: it placed third on the public EQ-Bench 4, which tests reading people in conversation, and larger story models write better long chapters. The best LLM for roleplay has more on that. The API docs list every field and error.
We have not published a tested SillyTavern set-up for it, so use Test Message before you rely on it.
Using Hemmingway Unlocked
Hemmingway Unlocked is the adults-only side of Hemmingway, with an uncensored model. Its API works in tools that take a custom OpenAI-style address, SillyTavern included.
- Address:
https://wise.hemmingway.io/v1 - Key: you create it inside Unlocked, in the account menu under API keys. A key needs a membership; the free trial is for the app.
- Model: pick the one the list shows once you are connected. A key made in Unlocked always reaches Unlocked's model.
It writes explicit and dark fiction between adults, and there is a short list of things it never writes, printed on the Unlocked page. Uncensored AI APIs says what to check before you build on one.
When the connection fails
- Nothing connects. Check the address ends in
/v1and has no/chat/completionson it. - A 401 error. The key is wrong, revoked or missing. On Hemmingway, a key that does not start with
hemmingway_live_is turned away this way. - A 402 error. On Hemmingway, the account's credit is used up, or the key draws on a plan that has ended.
- A 429 error. Too many requests are running at once, or a plan's allowance is used up. Wait and try again.
- The model list is empty. Type the model ID by hand.
Common questions
What API does SillyTavern use?
None of its own. SillyTavern is a front end that connects to a model you choose: a local backend such as KoboldCpp or Ollama, a hosted service such as OpenAI or OpenRouter, or any custom OpenAI-compatible endpoint.
Does SillyTavern need an API key?
Only for hosted services. Each service issues its own key, and you paste it into the API Connections panel. A model running on your own computer usually needs none.
Is the SillyTavern API free?
SillyTavern is free, and so is running a local model if you already have the hardware. Most hosted APIs are paid, and the price is set by the service, not by SillyTavern.
How do I add a custom API to SillyTavern?
Open API Connections, set the API type to Chat Completion, select Custom (OpenAI-compatible) as the source, and enter the address and key. Pick or type the model, then click Test Message. If it fails, add /v1 to the address.
Should I use Chat Completion or Text Completion?
Use Chat Completion for anything reached through an OpenAI-compatible address. Text Completion is what SillyTavern's guide to self-hosted models uses for a local backend such as KoboldCpp, where the prompt is sent as one long string.
Can I use Hemmingway with SillyTavern?
Hemmingway's API is in OpenAI's format, so add it as a custom endpoint with the address https://hemmingway.io/v1, a key from the API platform and the model hemmingway-27b. Hemmingway Unlocked, for adults, is at https://wise.hemmingway.io/v1.