An OpenAI-compatible API is an API, run by someone other than OpenAI, that accepts requests in the same format as OpenAI's Chat Completions and answers in the same format. Code and tools written for OpenAI can use it after three changes: the base URL, the API key and the model name. What the word does not promise is that everything else is the same. The model is a different model, some fields may be ignored, and the limits and errors are the provider's own.
We run one for our model, Hemmingway-1, so the examples here use it. The checks apply to any provider.
What "OpenAI-compatible" covers
At the least, five things.
- The address. A base URL, with the chat endpoint at
/chat/completionsbeneath it. - The key. Sent as a Bearer token in the
Authorizationheader. - The request. A JSON body with a
modeland a list ofmessages, each with a role and its content. - The reply. A chat completion object, with the text in
choices[0].message.contentand the token counts inusage. - Streaming. With
streamset to true, the reply arrives as server-sent events and ends withdata: [DONE].
If a service does those five the way OpenAI does, OpenAI's SDKs and the tools built on them can talk to it.
What changes in your code
| Setting | What you change it to |
|---|---|
| Base URL | The provider's address, usually ending in /v1 |
| API key | A key from the provider, not from OpenAI |
| Model name | The provider's own name for its model |
OpenAI's Python SDK takes the address as base_url when you create the client, or reads it from the OPENAI_BASE_URL environment variable. For Hemmingway the client becomes OpenAI(base_url="https://hemmingway.io/v1", api_key=...), and the call names the model as model="hemmingway-27b". A plain chat request needs nothing else changed.
Tools with no code in them work the same way. A front end such as SillyTavern has a custom OpenAI-compatible option where you enter the address and the key. The SillyTavern API guide shows where.
Who offers one
Hosted model APIs. A company runs the model and gives you a key. Hemmingway's is one.
Servers on your own machine. The common tools for running a model locally speak the same format, by their own docs:
- vLLM provides an HTTP server with the Chat Completions API.
- Ollama supports a subset of the OpenAI API at
http://localhost:11434/v1/. Its docs note that clients must send a key, and that Ollama ignores it. - LM Studio serves at
http://localhost:1234/v1, and its docs say existing OpenAI clients can be reused by changing the base URL. - KoboldCpp offers an OpenAI-compatible chat endpoint too.
The useful result is that your code does not care where the model lives. You can start on a hosted API and move to your own hardware later, or the other way round, by changing an address.
What may differ
This is where the work is. The format is shared, and the service behind it is not.
- Which endpoints exist. Chat completions is the common ground. Embeddings, legacy completions and the rest vary by provider. Hemmingway's API has chat completions and
GET /v1/models. - Which fields are honoured. Hemmingway passes a fixed list of fields to the model and drops any other. vLLM's docs say the
userfield is ignored. A field that is dropped rather than refused gives you no error, so a setting can quietly do nothing. Check the provider's list. - Extra fields. Providers add settings OpenAI does not have. Hemmingway accepts
top_k,min_p,repetition_penaltyandenable_thinking. In OpenAI's Python SDK these go inextra_body. - Thinking. Hemmingway-1 thinks before it answers. The thinking comes back in
reasoning_content, apart fromcontent.reasoning_efforttakeslow,mediumorxhigh, which is the default, andenable_thinkingset to false turns it off. - Token limits. On Hemmingway,
max_tokenscounts the thinking as well as the reply and is capped at 32,768, which is also the default. A low limit copied from other code can cut a reply short. - Pictures. Hemmingway reads up to four pictures a request, sent as
data:URLs. A picture given as a link is turned away with a 400. - Errors and limits. Status codes follow the usual pattern: 401 for a bad key, 402 when credit is used up, 429 when too many requests are running at once. The code names inside the error body are the provider's own, and so are the limits. Hemmingway allows 8 requests at once for an account's keys on credit, a body of up to 4 MB, and 13 minutes a request.
- The model name that comes back. Hemmingway has one model and names it
hemmingway-27bin every reply, whatever name you sent. - What is not there. In Hemmingway's case, web search and page reading are part of the app, not the API.
A checklist for switching
- Change the base URL, the key and the model name.
- Send one plain request and read the whole reply, not only the text.
- Turn on streaming and check your code handles the end of the stream.
- Test each extra you use, one at a time: function calling,
response_format, pictures. - Send a wrong key on purpose and see what your code does with the error.
- Set
max_tokenswith the provider's cap, and any thinking, in mind. - Run your real prompts through the old model and the new one, and compare the writing. Choosing an LLM API for writing has a way to do that blind.
Hemmingway's endpoint
| Base URL | https://hemmingway.io/v1 |
|---|---|
| Model | hemmingway-27b |
| Key | Made on the API platform. Keys start with hemmingway_live_ |
| Format | OpenAI's Chat Completions |
| Supports | Streaming, function calling, up to four pictures a request, control over thinking |
The API docs list every field, error and limit. Keys and prices are on the API platform.
Hemmingway-1's weights are open, so the same model can also run on your own hardware. Served with vLLM it speaks the same format on your machine, with Altworld/Hemmingway-1 as the model name. How to run it locally has the steps and the memory it needs.
For adults, Hemmingway Unlocked has its own endpoint at https://wise.hemmingway.io/v1. Uncensored AI APIs covers it.
Common questions
What does OpenAI-compatible mean?
It means an API accepts requests and returns replies in the same format as OpenAI's Chat Completions. Code written for OpenAI can use it by changing the base URL, the key and the model name.
Do I need an OpenAI account to use an OpenAI-compatible API?
No. The key comes from the provider you are calling, and nothing goes to OpenAI. You use OpenAI's format and, if you like, its SDK, but not its service.
Can I use the OpenAI SDK with another provider?
Yes. In the Python SDK, set base_url to the provider's address when you create the client, or set the OPENAI_BASE_URL environment variable. Then pass the provider's key and its model name.
Is an OpenAI-compatible API the same as OpenAI's API?
No. The request and reply formats match, but the model is different, and the provider decides which endpoints and fields it supports, what its limits are and what its errors say. Test the features you rely on.
Is Ollama OpenAI-compatible?
Partly. Ollama's docs say it supports a subset of the OpenAI API, at http://localhost:11434/v1/ on your own machine. A client has to send a key, and Ollama ignores it.
What is the base URL for the Hemmingway API?
It is https://hemmingway.io/v1, with the model hemmingway-27b and a key from the API platform. It takes OpenAI's Chat Completions format.