Hemmingway Try Hemmingway

Blog · Guide

OpenAI-compatible API

An OpenAI-compatible API takes requests in OpenAI's format, so your code needs a new base URL, key and model name. What stays the same and what may differ.

By the Hemmingway team ·

An OpenAI-compatible API is an API, run by someone other than OpenAI, that accepts requests in the same format as OpenAI's Chat Completions and answers in the same format. Code and tools written for OpenAI can use it after three changes: the base URL, the API key and the model name. What the word does not promise is that everything else is the same. The model is a different model, some fields may be ignored, and the limits and errors are the provider's own.

We run one for our model, Hemmingway-1, so the examples here use it. The checks apply to any provider.

What "OpenAI-compatible" covers

At the least, five things.

  • The address. A base URL, with the chat endpoint at /chat/completions beneath it.
  • The key. Sent as a Bearer token in the Authorization header.
  • The request. A JSON body with a model and a list of messages, each with a role and its content.
  • The reply. A chat completion object, with the text in choices[0].message.content and the token counts in usage.
  • Streaming. With stream set to true, the reply arrives as server-sent events and ends with data: [DONE].

If a service does those five the way OpenAI does, OpenAI's SDKs and the tools built on them can talk to it.

What changes in your code

SettingWhat you change it to
Base URLThe provider's address, usually ending in /v1
API keyA key from the provider, not from OpenAI
Model nameThe provider's own name for its model

OpenAI's Python SDK takes the address as base_url when you create the client, or reads it from the OPENAI_BASE_URL environment variable. For Hemmingway the client becomes OpenAI(base_url="https://hemmingway.io/v1", api_key=...), and the call names the model as model="hemmingway-27b". A plain chat request needs nothing else changed.

Tools with no code in them work the same way. A front end such as SillyTavern has a custom OpenAI-compatible option where you enter the address and the key. The SillyTavern API guide shows where.

Who offers one

Hosted model APIs. A company runs the model and gives you a key. Hemmingway's is one.

Servers on your own machine. The common tools for running a model locally speak the same format, by their own docs:

  • vLLM provides an HTTP server with the Chat Completions API.
  • Ollama supports a subset of the OpenAI API at http://localhost:11434/v1/. Its docs note that clients must send a key, and that Ollama ignores it.
  • LM Studio serves at http://localhost:1234/v1, and its docs say existing OpenAI clients can be reused by changing the base URL.
  • KoboldCpp offers an OpenAI-compatible chat endpoint too.

The useful result is that your code does not care where the model lives. You can start on a hosted API and move to your own hardware later, or the other way round, by changing an address.

What may differ

This is where the work is. The format is shared, and the service behind it is not.

  1. Which endpoints exist. Chat completions is the common ground. Embeddings, legacy completions and the rest vary by provider. Hemmingway's API has chat completions and GET /v1/models.
  2. Which fields are honoured. Hemmingway passes a fixed list of fields to the model and drops any other. vLLM's docs say the user field is ignored. A field that is dropped rather than refused gives you no error, so a setting can quietly do nothing. Check the provider's list.
  3. Extra fields. Providers add settings OpenAI does not have. Hemmingway accepts top_k, min_p, repetition_penalty and enable_thinking. In OpenAI's Python SDK these go in extra_body.
  4. Thinking. Hemmingway-1 thinks before it answers. The thinking comes back in reasoning_content, apart from content. reasoning_effort takes low, medium or xhigh, which is the default, and enable_thinking set to false turns it off.
  5. Token limits. On Hemmingway, max_tokens counts the thinking as well as the reply and is capped at 32,768, which is also the default. A low limit copied from other code can cut a reply short.
  6. Pictures. Hemmingway reads up to four pictures a request, sent as data: URLs. A picture given as a link is turned away with a 400.
  7. Errors and limits. Status codes follow the usual pattern: 401 for a bad key, 402 when credit is used up, 429 when too many requests are running at once. The code names inside the error body are the provider's own, and so are the limits. Hemmingway allows 8 requests at once for an account's keys on credit, a body of up to 4 MB, and 13 minutes a request.
  8. The model name that comes back. Hemmingway has one model and names it hemmingway-27b in every reply, whatever name you sent.
  9. What is not there. In Hemmingway's case, web search and page reading are part of the app, not the API.

A checklist for switching

  1. Change the base URL, the key and the model name.
  2. Send one plain request and read the whole reply, not only the text.
  3. Turn on streaming and check your code handles the end of the stream.
  4. Test each extra you use, one at a time: function calling, response_format, pictures.
  5. Send a wrong key on purpose and see what your code does with the error.
  6. Set max_tokens with the provider's cap, and any thinking, in mind.
  7. Run your real prompts through the old model and the new one, and compare the writing. Choosing an LLM API for writing has a way to do that blind.

Hemmingway's endpoint

Base URLhttps://hemmingway.io/v1
Modelhemmingway-27b
KeyMade on the API platform. Keys start with hemmingway_live_
FormatOpenAI's Chat Completions
SupportsStreaming, function calling, up to four pictures a request, control over thinking

The API docs list every field, error and limit. Keys and prices are on the API platform.

Hemmingway-1's weights are open, so the same model can also run on your own hardware. Served with vLLM it speaks the same format on your machine, with Altworld/Hemmingway-1 as the model name. How to run it locally has the steps and the memory it needs.

For adults, Hemmingway Unlocked has its own endpoint at https://wise.hemmingway.io/v1. Uncensored AI APIs covers it.

Get an API keyTry Hemmingway

Common questions

What does OpenAI-compatible mean?

It means an API accepts requests and returns replies in the same format as OpenAI's Chat Completions. Code written for OpenAI can use it by changing the base URL, the key and the model name.

Do I need an OpenAI account to use an OpenAI-compatible API?

No. The key comes from the provider you are calling, and nothing goes to OpenAI. You use OpenAI's format and, if you like, its SDK, but not its service.

Can I use the OpenAI SDK with another provider?

Yes. In the Python SDK, set base_url to the provider's address when you create the client, or set the OPENAI_BASE_URL environment variable. Then pass the provider's key and its model name.

Is an OpenAI-compatible API the same as OpenAI's API?

No. The request and reply formats match, but the model is different, and the provider decides which endpoints and fields it supports, what its limits are and what its errors say. Test the features you rely on.

Is Ollama OpenAI-compatible?

Partly. Ollama's docs say it supports a subset of the OpenAI API, at http://localhost:11434/v1/ on your own machine. A client has to send a key, and Ollama ignores it.

What is the base URL for the Hemmingway API?

It is https://hemmingway.io/v1, with the model hemmingway-27b and a key from the API platform. It takes OpenAI's Chat Completions format.

Read next

  • Choosing the best LLM API for writingWhat to look for in an LLM API for writing features: human-sounding output, no commentary around the message, context, OpenAI compatibility and open weights.
  • SillyTavern APIWhat an API connection is in SillyTavern, the kinds you can use, where an API key goes, and how to add a custom OpenAI-compatible endpoint such as Hemmingway.
  • How to run Hemmingway-1 locallyRun Hemmingway-1 locally with vLLM or Transformers: the commands, how much memory a 27B model needs, and the hosted API if you have no GPU.
  • What is Hemmingway AI?What is Hemmingway AI? A plain explanation of the lab, its open 27B model Hemmingway-1, the apps for Mac, Windows, Android and web, and the API.
  • The best local LLM for writingThe best local LLM for writing is the largest writing model your memory can hold. A sizing rule, the tools that run one, a test, and where Hemmingway-1 fits.
  • Uncensored AI APIAn uncensored AI API serves a model that writes adult and dark content without refusing. The three ways to get one, what to check, and Hemmingway Unlocked's.