Hemmingway Try Hemmingway

Blog · Guide

Uncensored LLM

An uncensored LLM writes the adult and dark fiction ordinary models refuse. How models are made that way, how to run one locally, and how to use one by API.

By the Hemmingway team ·

An uncensored LLM is a language model that writes what ordinary assistants refuse, above all adult and dark fiction, without stopping to say no, lecture you or soften the scene. There are three routes to one: a model is trained further so that it answers where the original refused, its refusal behaviour is removed after training (abliteration), or a company serves it under rules that allow adult content. You can run one on your own computer, or use a hosted one through an app or an API.

It is for adults, and it never means a model with nothing it will not do. A model on your own machine is still bound by the law and by its licence, and a hosted one by its service's terms. We make Hemmingway Unlocked, a hosted uncensored model for adults, so this guide says where it fits and where a local model suits you better.

Why ordinary models refuse

Two separate things produce a refusal.

  • The model. In training, a model is taught to turn certain requests down. That habit sits in its weights and goes wherever the model runs.
  • The service around it. A hosted assistant can also check what goes in and what comes out, and stop a reply the model would have written.

General assistants serve a general audience, so both are set cautiously. On 27 September 2026 we sent the same adult roleplay request to ChatGPT, Claude and Gemini, and all three refused. Why ChatGPT refuses roleplay has the test and the reasons.

What makes a model uncensored

Fine-tuning. Someone takes an open model and trains it further, so that it answers where the original said no. The Dolphin series from Cognitive Computations is one example. The model card for Dolphin 3.0, built on Llama 3.1 8B, says the model does not impose its own ethics or guidelines, and that the person running it decides them.

Abliteration. This removes refusals without any retraining. Maxime Labonne's article on Hugging Face explains the idea: refusal in a model is carried by one direction in its internal activations, and if the model is stopped from representing that direction, it loses the ability to refuse. His article also reports the cost. The abliterated model in his example lost quality, and he trained it further to repair it. One published example is mlabonne/gemma-3-27b-it-abliterated, made from Google's Gemma 3 27B. Its card calls the result fairly experimental and says Gemma 3 resisted the technique more than other models did.

Hosting under different rules. A company serves a model behind its own app and API, and publishes what it will and will not write. You need no hardware, and you are trusting its rules and its privacy promises. Hemmingway Unlocked is one of these.

What uncensored does not mean

  • Better writing. Taking refusals out does not improve the prose, and by Labonne's account it can make a model worse. Plenty of unfiltered models write flat, repetitive scenes.
  • No law and no licence. What is illegal to write is illegal with any model. Open models also come with licences, and they differ: the two cards above list the Llama 3.1 licence and the Gemma licence. Read the card before you build anything on a model.
  • Private. A local model is private because nothing leaves your machine. A hosted one is only as private as its terms say.

Running an uncensored LLM locally

You need three things.

  1. A runner. Ollama, LM Studio and KoboldCpp are the usual ones. KoboldCpp's wiki describes it as text-generation software for GGML and GGUF model files. LM Studio has a Developer tab where you start a local server. Each of the three can offer an OpenAI-compatible endpoint on your own machine, so a front end such as SillyTavern can connect to it. The SillyTavern API guide covers that part.
  2. A model file. Models are downloaded from Hugging Face. Read the model card first: what it was built on, how the refusals were removed, the licence, and any warning from the person who made it.
  3. Enough memory. At 16-bit precision a parameter takes 2 bytes, so a 27B model needs about 54 GB for the weights alone, before any room for context. On an ordinary computer that means a smaller model, or a quantised file, which stores each parameter in fewer bits.

People ask which is the best uncensored LLM to run locally. We have not tested local uncensored models, so we do not rank them here. A fair way to choose:

  • Pick the largest model that fits in your memory with room left for the conversation.
  • Check how it was made. An abliterated model can lose some quality, so see whether it was trained further afterwards.
  • Try it on a scene you know well, and judge the writing, not only whether it refused.

SillyTavern's own docs state the trade-off plainly: local models are free to use and have no content filter, and most are not as strong as the hosted ones.

Local or hosted

Local modelHosted uncensored model
PrivacyNothing leaves your computerMessages go to the service to be answered
HardwareYours, and it caps the size of the modelNone
WritingLimited by what fits in your memoryWhatever the service runs
RulesThe law and the model's licenceThe service's terms, which you can read first
CostNo subscription, once you own the hardwareUsually paid after a trial

An uncensored LLM API

If you would rather call a model than host one, look for an API in OpenAI's format, because most front ends and SDKs already speak it. Hemmingway Unlocked has one at https://wise.hemmingway.io/v1. You create a key inside Unlocked, in the account menu under API keys. A key needs a membership; the free trial is for the app. The guide to uncensored AI APIs has the details and what to check before you build on one.

Where Hemmingway Unlocked fits

Hemmingway Unlocked is the adults-only side of Hemmingway: the same app, with an uncensored model, at wise.hemmingway.io. It asks first whether you are 18 or over, then shows its rules.

  • What it writes. Sex between adults, explicit and in full, and dark, violent stories, without refusing or lecturing on grounds of taste.
  • What it never writes. Nothing sexual involving anyone under 18, a real person, animals, family or anyone who does not consent. And never torture, help hurting or harassing a real person, weapons, making drugs, harmful software, or help with self-harm. The list is in our terms, and every answer has a Report button.
  • Where chats are kept. In your browser, on your device. They are not saved to your account or our servers, and they are never trained on. Each message passes through our servers to be answered and is not kept.
  • What it costs. It is free to try, with no card. After the trial it is a monthly membership.

Open Hemmingway UnlockedTry Hemmingway

It does not fit every case. If nothing may leave your computer, run a local model. Unlocked's weights are not out yet. We plan to open-source it soon. Hemmingway-1, our regular model, has open weights under CC BY-NC 4.0, but it is not the uncensored one.

Common questions

What is an uncensored LLM?

It is a language model that writes adult and dark content that ordinary assistants refuse. Some are open models changed by fine-tuning or abliteration and run on your own computer. Others are hosted by a service whose rules allow adult content. None is free of rules: the law, the model's licence and the service's terms still apply.

What is the best uncensored LLM to run locally?

We have not tested local uncensored models, so we do not name a winner. Choose the largest one that fits your memory with room for the conversation, read its model card for how it was made and its licence, and try it on a scene you know well.

What does abliterated mean?

An abliterated model has had its refusal behaviour removed without retraining. The method finds the direction inside the model that carries refusal and stops the model from using it. It can cost some writing quality, which is why some abliterated models are trained further afterwards.

Is there an uncensored LLM API?

Yes. Hemmingway Unlocked has an OpenAI-compatible API at https://wise.hemmingway.io/v1, with a key you create in Unlocked. A key needs a membership. You can also run an open model yourself and call it on your own machine.

Can I run the Hemmingway Unlocked model locally?

Not yet. Its weights are not published, and we plan to open-source it soon. Hemmingway-1, our regular model, has open weights, but it is not the uncensored one.

Read next

  • Uncensored AIUncensored AI does not refuse or lecture adults over dark or sexual writing. What the word means, the three ways to get one, and where the limits still sit.
  • Uncensored AI APIAn uncensored AI API serves a model that writes adult and dark content without refusing. The three ways to get one, what to check, and Hemmingway Unlocked's.
  • Open-source models for roleplayWhy people run open-source models for roleplay, how to set one up with vLLM or an OpenAI-compatible API, and how Hemmingway-1 compares for character chat.
  • Uncensored AI roleplayUncensored AI roleplay lets adults play dark, violent or sexual scenes without the story being stopped. What it means, what to look for, and how to set one up.
  • Venice AI alternativesVenice AI alternatives, by what you came for: privacy, fewer refusals, or images and characters in one place. What each option keeps and what it allows.
  • The best LLM for roleplayThe best LLM for roleplay stays in character, keeps the story in context and does not refuse your scenes. How to judge each, local or API, with a test to run.