Hemmingway Try Hemmingway

Blog · Guide

The best local LLM for writing

The best local LLM for writing is the largest writing model your memory can hold. A sizing rule, the tools that run one, a test, and where Hemmingway-1 fits.

By the Hemmingway team ·

The best local LLM for writing is the largest model built for writing that fits in your machine's memory with room left for context. Memory decides first, so start there: at 16-bit precision a model needs about 2 bytes for every parameter, which puts a 27B model at roughly 54 GB for the weights alone. Hemmingway-1, our open-weights writing model, is that size. On a smaller machine you would run a smaller open model or a quantised one. We have not tested those for writing, so we do not rank them here.

What follows is a sizing rule, the two families of tools that run a model locally, a short test for writing quality, and an honest account of where Hemmingway-1 fits.

Start with memory

A model has to sit in memory to run, and preferably in a graphics card's memory, where it is fast. The sum is simple.

  • Weights. Parameters times bytes per parameter. At 16-bit precision that is 2 bytes each, so take the parameters in billions and double them to get gigabytes. For 27B, about 54 GB.
  • Context. Every token the model is holding takes memory on top of the weights. The longer the context you allow, the more you need.
  • Quantised files. A quantised model stores each parameter in fewer bits, so the same sum comes out smaller. It costs some quality, and how much depends on the model. We have not tested quantised versions of any model, our own included.
  • More than one card. If one graphics card is too small, the model can be spread over several.
  • Ordinary memory. Part of a model can be put in the computer's ordinary memory. It works and it is slow: fine for a test, not for daily use.

In general terms, that sorts machines like this.

Your machineWhat it can hold
A laptop with no separate graphics cardA small model, quantised, running from ordinary memory. It runs, slowly
A desktop with one consumer graphics cardA small or mid-sized model, usually quantised
A large data-centre GPU, or several cardsA large model at full precision, such as a 27B one

We have not published tested hardware set-ups, so treat this as a sizing rule and not as advice to buy a particular card.

Two families of tools

Full weights: vLLM and Transformers. These load a model as its makers published it. This is how Hemmingway-1 runs. With vLLM it is one command:

vllm serve Altworld/Hemmingway-1 --max-model-len 262144

That starts a server on your own machine that speaks OpenAI's format, so any writing tool that accepts a custom address can use it. A smaller --max-model-len sets aside less memory for context. How to run Hemmingway-1 locally has the full steps.

Quantised files: Ollama, LM Studio and KoboldCpp. These are the usual way to run a model on an ordinary computer. KoboldCpp's wiki describes it as text-generation software for GGML and GGUF model files. LM Studio has a Developer tab where you start a local server. SillyTavern's docs call Ollama the easiest of the llama.cpp-based options to set up, with a catalogue of models to download. All three can also offer an OpenAI-compatible endpoint on your own machine.

There is no official GGUF or quantised release of Hemmingway-1. The only weights we publish are the full ones, so it belongs to the first family.

What to ask of a local writing model

A model can be clever and still write badly. Before you settle on one, give it four small jobs and read the results as the person who would receive them.

A hard short message.

"Write the message I send a friend to say I can't lend him money again."

Look for the message itself. Not a preamble, not three versions, not a note on tone afterwards.

Notes into a paragraph.

"Turn these notes into one paragraph for my team: shipment late, supplier's fault, new date Thursday, nothing needed from them."

Look for every fact you gave and none you did not.

Your own voice. Paste three messages you really sent and ask for a fourth in the same voice. Look for your rhythm and your words, not a tidier stranger's.

A story opening.

"Write the first 200 words of a story about a lighthouse keeper who finds a second set of footprints."

Look for one specific detail you did not expect, and no moral.

Then count. How many of the four could you send or keep without editing? The two failures to watch for are a message wrapped in commentary and a fact you never gave. They matter more than any score.

Where Hemmingway-1 fits

Hemmingway-1 is a 27B model built on Qwen3.8-27B and trained further for one job: writing that sounds like a person wrote it. Its weights are on Hugging Face at Altworld/Hemmingway-1 under CC BY-NC 4.0, free for non-commercial use, with commercial use by agreement. Its context is 262,144 tokens.

In our own blind tests it came first on everyday writing, CommunicationBench, with 1026, level with Fable 5.1, and first on Human-Likeness with 1032. The base model it started from scored 954 and 952 on the same two tests, which shows what the training changed. Those benchmarks are ours. The scores are on the model card, and the best open-source LLM for writing goes through them.

Where it falls short as a local model:

  • It is large. Plan for more than 54 GB of GPU memory. That is a data-centre card or several smaller ones, not a laptop.
  • Long fiction. Larger story models write better novels and long chapters.
  • English first. Test it yourself in any other language.
  • Not for decisions. It can be wrong and still sound certain, so do not use it for medical, legal or money questions.

If your machine is too small for it

You have three choices.

  1. A smaller open model, locally. Several families come in small sizes. Google's Gemma 3, for one, is listed on its model card in 1B, 4B, 12B and 27B sizes, with open weights under the Gemma licence. We have not tested it for writing. Whatever you choose, read the licence on its card, and run the four jobs above.
  2. Hemmingway-1 over the API. The same weights are served at https://hemmingway.io/v1 as hemmingway-27b, in OpenAI's format. Your text goes to our server to be answered, and the API platform states that nothing you send and nothing the model writes is logged. Keys and prices are on the API platform. What OpenAI-compatible means covers the switch.
  3. The app. If you want to write with it and not build on it, the Hemmingway app uses the same model.

Running a model yourself makes sense when your text must stay on your own machine, when you want to fine-tune, or when you already own the hardware. Otherwise the API is less work.

Get an API keyTry Hemmingway

Common questions

What is the best local LLM for writing?

The largest model built for writing that fits in your memory. Hemmingway-1 is an open-weights model trained for writing that sounds like a person, and it came first on our own everyday-writing test, but it needs more than 54 GB of GPU memory. On smaller machines, run a smaller open model and test it on your own requests.

How much memory does a local LLM need?

At 16-bit precision, about 2 bytes per parameter, so double the parameters in billions to get gigabytes: roughly 54 GB for a 27B model. The context needs memory on top. A quantised file is smaller, at some cost in quality.

Can I run a writing LLM on a laptop?

A small, quantised one, yes, and it will be slow without a separate graphics card. A 27B model at full precision needs well over 54 GB of GPU memory, which most laptops do not have. For a large model on a laptop, the practical route is an API or an app.

Can I run Hemmingway-1 in Ollama or LM Studio?

There is no official GGUF or quantised release of Hemmingway-1, and we have not tested any other version. The weights we publish are the full ones, which run with vLLM or Transformers.

Is a local LLM private?

Yes, in the plain sense that nothing you write leaves your computer once the model is downloaded. That is the main reason to run one locally.

Read next

  • How to run Hemmingway-1 locallyRun Hemmingway-1 locally with vLLM or Transformers: the commands, how much memory a 27B model needs, and the hosted API if you have no GPU.
  • The best open-source LLM for writingHemmingway-1 is a 27B open-weights model under CC BY-NC 4.0, built for writing that sounds like a person. How it scores, what it needs, and how to run it.
  • Choosing the best LLM API for writingWhat to look for in an LLM API for writing features: human-sounding output, no commentary around the message, context, OpenAI compatibility and open weights.
  • OpenAI-compatible APIAn OpenAI-compatible API takes requests in OpenAI's format, so your code needs a new base URL, key and model name. What stays the same and what may differ.