Hemmingway Try Hemmingway

Blog · Guide

SillyTavern API

What an API connection is in SillyTavern, the kinds you can use, where an API key goes, and how to add a custom OpenAI-compatible endpoint such as Hemmingway.

By the Hemmingway team ·

SillyTavern has no model of its own. It is a front end: it builds a prompt from your chat, your character and your settings, and sends it to a model through an API connection that you set up. That connection can be a model running on your own computer, a hosted service you have a key for, or any endpoint that speaks OpenAI's format. You choose it in the API Connections panel, which is the second button in the top bar.

This guide covers what the choices mean, where an API key goes, and how to add a custom OpenAI-compatible endpoint. We make Hemmingway, so the last part shows its two addresses. The steps are the same for any compatible API. What we say about SillyTavern's menus comes from its own docs at docs.sillytavern.app.

What an API connection is

Two settings decide how SillyTavern talks to a model.

  • The API type. SillyTavern's docs describe two main ones. Chat Completion sends the prompt as a list of messages between the user, the assistant and the system. Text Completion turns the whole prompt into one long string, and the model continues it. With Text Completion you can apply an instruct template, which helps the model respond to instructions.
  • The source or backend. This is who answers: a program on your machine, or a service on the internet.

An OpenAI-compatible address is always a Chat Completion connection.

The kinds of API

KindExamples in SillyTavern's docsGood forWatch out for
Local backendsKoboldCpp, llama.cpp, Ollama, Oobabooga's TextGeneration WebUI, TabbyAPIPrivacy, and no content filterYou download the models and need the hardware to run them
Hosted APIsOpenAI, Claude, Google AI Studio, Mistral, OpenRouter, DeepSeek, NovelAIStronger models with nothing to installContent filtering of varying degrees, and most are paid
Custom OpenAI-compatibleLM Studio, LiteLLM, LocalAI, or any service with a compatible addressAnything not on the built-in listSillyTavern does not promise every endpoint works

The docs sum up the trade-off in much the same words: local models are free to use and unfiltered, and most are not as strong as the hosted ones. Hosted models need nothing from your computer and are stronger, and all of them filter content to some degree.

API keys

A key is how a hosted service knows a request is yours and which account to charge.

  • Where it comes from. Each service issues its own, on its own site. SillyTavern does not give you one.
  • Where it goes. In the API Connections panel, in the key field for the source you picked.
  • Local backends. A program on your own machine usually needs no key. Ollama's docs say clients must send a key value, but Ollama ignores it.
  • Keep it private. A key is as good as a password for spending your credit. Do not paste it into a shared preset, a screenshot or a forum post.

If you switch between models often, SillyTavern's Connection Profiles save the API type, the model, the server address and the key together, so changing model is one choice in a list. Set the connection up, click Create, and name it.

How to add a custom OpenAI-compatible endpoint

These are the steps from SillyTavern's docs.

  1. Open API Connections, the second button in the top bar.
  2. Set the API type to Chat Completion.
  3. For Chat Completion Source, select Custom (OpenAI-compatible).
  4. Enter the endpoint address, and the API key if the service needs one.
  5. Choose the model. If the endpoint lists its models, you pick from a dropdown. If it does not, type the model ID into the text field.
  6. Click Test Message. It sends a short prompt to check the connection.

Two details save most of the trouble. If it will not connect, add /v1 to the end of the address. And do not add /chat/completions to it.

There is also an option called Bypass API status check, which stops SillyTavern warning you that the endpoint is not working. And under Prompt Post-Processing you can have SillyTavern reshape the messages before sending: merge consecutive messages from the same role, allow only one system message, require a user message first, or fold everything into a single user message. Leave it at None unless an endpoint complains about the order of messages.

SillyTavern's docs are clear that they give no support for custom endpoints and no promise that every one works. Test before a long session.

Connecting a local backend

For KoboldCpp, SillyTavern's guide to self-hosted models gives these settings: API set to Text Completion, API Type set to KoboldCpp, the server address http://127.0.0.1:5001/, then Connect.

LM Studio and Ollama can also be reached as custom OpenAI-compatible endpoints. In LM Studio you start the server in the Developer tab, and its docs use http://localhost:1234/v1 as the address. Ollama's docs give http://localhost:11434/v1/.

For how much memory a model needs, and how to serve Hemmingway-1's open weights yourself, see open-source models for roleplay.

Using Hemmingway in SillyTavern

Hemmingway's API takes OpenAI's Chat Completions body, so it goes in as a custom endpoint.

  • Address: https://hemmingway.io/v1
  • Key: one you make on the API platform, where the prices are too. Keys start with hemmingway_live_.
  • Model: hemmingway-27b

Three things to know. The model thinks before it answers, and the thinking arrives in its own field, reasoning_content, apart from the reply. The reply limit counts the thinking too, so do not set it very low. And the model was built for talking to people and for writing, not as a roleplay model: it placed third on the public EQ-Bench 4, which tests reading people in conversation, and larger story models write better long chapters. The best LLM for roleplay has more on that. The API docs list every field and error.

We have not published a tested SillyTavern set-up for it, so use Test Message before you rely on it.

Get an API keyTry Hemmingway

Using Hemmingway Unlocked

Hemmingway Unlocked is the adults-only side of Hemmingway, with an uncensored model. Its API works in tools that take a custom OpenAI-style address, SillyTavern included.

  • Address: https://wise.hemmingway.io/v1
  • Key: you create it inside Unlocked, in the account menu under API keys. A key needs a membership; the free trial is for the app.
  • Model: pick the one the list shows once you are connected. A key made in Unlocked always reaches Unlocked's model.

It writes explicit and dark fiction between adults, and there is a short list of things it never writes, printed on the Unlocked page. Uncensored AI APIs says what to check before you build on one.

When the connection fails

  • Nothing connects. Check the address ends in /v1 and has no /chat/completions on it.
  • A 401 error. The key is wrong, revoked or missing. On Hemmingway, a key that does not start with hemmingway_live_ is turned away this way.
  • A 402 error. On Hemmingway, the account's credit is used up, or the key draws on a plan that has ended.
  • A 429 error. Too many requests are running at once, or a plan's allowance is used up. Wait and try again.
  • The model list is empty. Type the model ID by hand.

Common questions

What API does SillyTavern use?

None of its own. SillyTavern is a front end that connects to a model you choose: a local backend such as KoboldCpp or Ollama, a hosted service such as OpenAI or OpenRouter, or any custom OpenAI-compatible endpoint.

Does SillyTavern need an API key?

Only for hosted services. Each service issues its own key, and you paste it into the API Connections panel. A model running on your own computer usually needs none.

Is the SillyTavern API free?

SillyTavern is free, and so is running a local model if you already have the hardware. Most hosted APIs are paid, and the price is set by the service, not by SillyTavern.

How do I add a custom API to SillyTavern?

Open API Connections, set the API type to Chat Completion, select Custom (OpenAI-compatible) as the source, and enter the address and key. Pick or type the model, then click Test Message. If it fails, add /v1 to the address.

Should I use Chat Completion or Text Completion?

Use Chat Completion for anything reached through an OpenAI-compatible address. Text Completion is what SillyTavern's guide to self-hosted models uses for a local backend such as KoboldCpp, where the prompt is sent as one long string.

Can I use Hemmingway with SillyTavern?

Hemmingway's API is in OpenAI's format, so add it as a custom endpoint with the address https://hemmingway.io/v1, a key from the API platform and the model hemmingway-27b. Hemmingway Unlocked, for adults, is at https://wise.hemmingway.io/v1.

Read next

  • Open-source models for roleplayWhy people run open-source models for roleplay, how to set one up with vLLM or an OpenAI-compatible API, and how Hemmingway-1 compares for character chat.
  • OpenAI-compatible APIAn OpenAI-compatible API takes requests in OpenAI's format, so your code needs a new base URL, key and model name. What stays the same and what may differ.
  • Uncensored AI APIAn uncensored AI API serves a model that writes adult and dark content without refusing. The three ways to get one, what to check, and Hemmingway Unlocked's.
  • The best LLM for roleplayThe best LLM for roleplay stays in character, keeps the story in context and does not refuse your scenes. How to judge each, local or API, with a test to run.
  • Janitor AI alternativesJanitor AI alternatives by kind: other character sites, a front end with your own model, hosted uncensored chat and Hemmingway Unlocked. What each suits.