An unfinished manuscript is the most sensitive document most writers own. The publishers are not on the cc list; the agents are not in the chat; the characters are not real people. Yet every keystroke a writer puts into a hosted AI tool gets sent to a vendor’s data centre, where it sits until the vendor’s retention policy says otherwise. Three hosted services dominate the fiction-writing niche today: NovelAI, Novelcrafter, and the general-purpose assistants that writers quietly paste chapters into anyway. The alternative that does not involve trusting any of them is the same one covered across this site’s local-AI cluster: a model on your hardware, a front end that knows how to talk to it, and the discipline to never send the manuscript anywhere by default.
This page walks through what “local AI for creative writing” actually looks like in October 2026, what the models and front ends are, and which combinations of model size and front end cover the four jobs most novelists, screenwriters, and game writers actually need - drafting prose, worldbuilding, dialogue, and editing passes.
TL;DR
- Local creative-writing AI is two pieces, not one product. A model runner (Ollama or llama.cpp) supplies the weights and inference. A front end (SillyTavern, KoboldCpp, text-generation-webui, Open WebUI) supplies the prose workflow - system prompts, world info, lorebooks, character cards, sampler knobs.
- The mainstream chat models cover drafting well, but creative-writing fine-tunes cover prose voice better. TheDrummer’s Cydonia-24B-v4.3 (Apache-family Mistral lineage) and Skyfall-31B-v4.2 are prominent creative-writing fine-tunes on Hugging Face, and the model’s own model card reads: “Foundational models today are optimized for non-creative uses, and I believe there is a place for AI in creativity and entertainment.”
- The hosted alternatives are not training on your prose by default - they just see it. NovelAI’s homepage positions the story generator as a writing space built around the way the writer creates, with subscriptions starting at $10/month. Novelcrafter starts at $4/m with a 21-day free trial and is not locked in to a single model provider - it can route prompts to outside AI services.
- If you want zero bytes of your manuscript leaving your machine, the path is Ollama or llama.cpp on your host, SillyTavern or KoboldCpp on top of it, and a model with a permissive licence. Every option below is documented with primary sources.
What “local creative writing” actually means
Four jobs cover most AI-assisted fiction work:
- Drafting prose. Produce text in a writer’s voice given a scene, a character, a point of view. The model reads the prior scene and writes the next few paragraphs.
- Worldbuilding. Answer “tell me what the magic system says about debt” or “describe a market in a port city” and have the response respect the lorebook already set.
- Dialogue. Have the model stay in character over many turns, with consistent voice and continuity.
- Editing passes. Rewrite for clarity, tighten a paragraph, or push the prose toward a target style.
A “local” deployment has to handle all four without sending the manuscript to a third-party LLM. The model has to run on hardware you control. The front end has to store the lorebook, the system prompt, and the chat history on the same hardware. Anything that bounces through a vendor API - even if the model is open-weight - breaks the chain.
Two pieces have to come together for this to work:
- A model runner that keeps weights and inference on-device. Ollama is the most common pick; llama.cpp is the underlying engine that KoboldCpp and text-generation-webui wrap. Ollama’s library has the broadest catalogue for the dense model sizes a writer is likely to fit.
- A front end that knows how to handle prose workflows. General chat UIs work for short tasks, but long-form fiction wants character cards, lorebooks, world info, sampler controls, and the ability to branch a scene. That is where the creative-writing front ends separate from the general ones.
Models worth pulling
The mainstream chat models cover drafting. The creative-writing fine-tunes cover prose voice, characterisation, and the willingness to keep going without refusals. The Ollama tags below are pulled from the live library pages today.
| Job | Model | Ollama tag (size) | Context | Licence |
|---|---|---|---|---|
| 8 GB laptop, drafting / dialogue | qwen3:4b | 2.5 GB | 256K | see model card |
| 16 GB tier, long context | qwen3:8b | 5.2 GB | 40K | see model card |
| 16 GB tier, prose voice | mistral-small3.2:24b | 15 GB | 128K | see model card |
| 24 GB tier, prose + editing | qwen3:14b | 9.3 GB | 40K | see model card |
| 24 GB tier, frontier dense | mistral-small:24b | 14 GB | 32K | Apache 2.0 (Mistral Small 3 family) |
| 24 GB tier, MoE prose | qwen3:30b | 19 GB | 256K | see model card |
| 64 GB+ workstation | llama3.3:70b | 43 GB | 128K | Llama 3.3 Community License |
| Frontier MoE (multi-GPU) | qwen3:235b | 142 GB | 256K | see model card |
The Qwen3 family, per the Ollama qwen3 page, is “the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models.” Mistral Small 3.2 is a minor update of Small 3.1; per the Ollama mistral-small3.2 page, it “improves on function calling, instruction following, and less repetition errors.” Mistral Small 3 itself is “Apache 2.0 License” with a “32k context window,” per the Ollama mistral-small page.
These mainstream tags are not creative-writing-specific. They will draft; they will refuse on content many writers want to write; and the prose voice tends toward “general-purpose assistant.” That is where the fine-tunes come in.
Creative-writing fine-tunes
For prose voice and characterisation, the Hugging Face ecosystem is the place to look. TheDrummer’s “Portfolio 2026” collection is a prominent creative-writing fine-tune family on the platform, and the Cydonia-24B-v4.3 model card and Skyfall-31B-v4.2 model card make the position clear:
“Foundational models today are optimized for non-creative uses, and I believe there is a place for AI in creativity and entertainment.”
TheDrummer’s stated evaluation criteria are Writing, Dynamism, Imagination, Attitude, Morality, Formatting, Adherence, Knowledge, and Perception - a vocabulary borrowed from roleplay communities, not from the leaderboard-style evaluation suites that drive most model releases. Both Cydonia and Skyfall are derived from Mistral-Small-3.x, which is Apache 2.0. Cydonia’s model card states the base lineage explicitly: “mistralai/Mistral-Small-3.1-24B-Base-2503 -> mistralai/Mistral-Small-3.2-24B-Instruct-2506 -> Cydonia-24B-v4.3 (this model).” Skyfall’s lineage is “mistralai/Mistral-Small-3.1-24B-Base-2503, with subsequent finetunes through Mistral-Small-3.2-24B-Instruct-2506 and Magistral-Small-2509.”
For Ollama users, these arrive as GGUF quantisations in community repos. For llama.cpp / text-generation-webui / SillyTavern users, the Hugging Face GGUF is downloaded directly. The base model for both is Mistral-Small-3.x (Apache 2.0 per the Ollama mistral-small page), but the Hugging Face model pages I read do not restate the licence for the fine-tune itself, and writers should confirm the fine-tune author’s redistribution terms before publishing anything commercially that includes generated prose from a fine-tune - the licence chain matters.
Front ends for the writing workflow
The model alone is not enough. The front end is where the system prompt, the lorebook, the character card, and the sampler controls live.
SillyTavern. Per the SillyTavern docs, it is “[a] locally installed user interface that allows you to interact with text generation LLMs, image generation engines, and TTS voice models,” AGPL-3.0 licensed. The features that matter for fiction:
- Character cards - “A character card is a collection of prompts that set the behavior of the LLM.” A card can be a fictional character, an abstract scenario, or “an assistant tailored for a specific task.”
- World Info support - “create rich lore or save tokens on your character card.”
- Built-in RAG support - “add documents to your chats for the AI to reference,” useful for pinning a style guide or a world bible.
- Stable Diffusion / FLUX / DALL-E image generation - cover art and character portraits, all from the same UI.
- Auto-Summary of the chat history - the only way to keep a 100K-token roleplay readable.
SillyTavern is the front end most creative-writing users land on first. It connects to OpenAI-compatible APIs, KoboldAI, Tabby, AI Horde, and others, so it can talk to a local Ollama endpoint or to a hosted model with the same interface.
KoboldCpp. Per the KoboldCpp GitHub, it is “free and open-source software for running GGUF large language models (LLMs) on your own computer,” “One File. Zero Install,” and “Inspired by KoboldAI and built on llama.cpp,” AGPL-3.0. Its modes include “chat, adventure, instruct, or story writing modes,” and the recommended example models are Qwen3-VL-8B, L3-8B-Stheno-v3.2, and Gemma3-4B. The story-writing mode is the closest thing in the local-AI ecosystem to a purpose-built fiction interface, and the single-file portability (no installer, just an executable) makes it the easiest to set up on a writer’s existing laptop.
text-generation-webui (oobabooga / “TextGen”). Per the text-generation-webui GitHub, it is now rebranded as “TextGen” and described as an open-source desktop app for running LLMs locally, AGPL-3.0 licensed. Modes relevant to fiction: “instruct, chat-instruct, and chat modes with auto-formatted Jinja2 prompts; message editing, branching; notebook tab.” The notebook tab is the key fiction feature - it lets a writer generate free-form text without a chat thread, then copy the best continuation into the manuscript. Backends include llama.cpp, ik_llama.cpp, Transformers, ExLlamaV3, and TensorRT-LLM.
Open WebUI. Per the Open WebUI GitHub, it is self-hosted and supports any OpenAI-compatible endpoint (so an Ollama backend works), with Local RAG over uploaded documents. Open WebUI is the most general of the four - it is the right pick if the same writer also wants chat with their notes, a RAG pipeline over a research folder, and the ability to switch models without learning a new UI. It does not have character cards or world info in the SillyTavern sense.
The hosted alternatives
Three hosted services writers compare local setups against.
NovelAI. Per the NovelAI homepage, the story generator is “Powered by GLM 4.6” and is positioned as “A writing space built around the way you create, not the other way around.” Subscriptions “start at $10/month, perfect for casual creators with access to both our AI anime image generator and text generator.” Image generation is “NovelAI Diffusion V5.” The privacy promise: “Prompts are encrypted,” “Images are not stored on servers (only during sessions unless saved),” “Complete anonymity offered,” and “No restrictions on commercial use of generated images.” What it does not say: that the manuscript text is not seen by the model. Every generation request ships the prior chapter to NovelAI’s servers.
Novelcrafter. Per the Novelcrafter homepage, it is a cloud-hosted writing tool with a 220k-author community, plans “starting from 4 USD/m,” a “21-day free trial,” and “No credit card required.” The Codex is a wiki/story-bible/world builder that “automatically keeps track and links” characters, places, and lore, and can be shared across books in a series. The platform is “Not locked in, ever” - users can connect to outside AI providers. The catch is the same as NovelAI’s: until the user connects a local model through that outside-provider slot, every chapter is hosted.
General-purpose assistants. ChatGPT, Claude, and Gemini are the AI tools most writers paste chapters into for editing or “what would you do with this paragraph” feedback. Each has published data-retention policies and a way to turn off training, but each also receives the chapter by default.
The privacy delta is the only honest argument for the local stack. Quality is mixed - a frontier hosted model still writes more careful prose than a 24B local model - but the bytes of the manuscript leave the machine, and that is what changes when the manuscript moves to a local model.
What the local stack does not solve
Three honest limitations.
- Frontier-model quality on small hardware. A 4B model will draft serviceable prose; it will not match a Claude Opus 5.5 or a GPT-6 for the subtle paragraph-level work a careful editor does. The local models close the gap with the right system prompt and a 30B+ quant, but the gap exists.
- Lorebook scale. A 200-entry world bible eats context. SillyTavern’s World Info entries are vector-matched, which helps; Open WebUI’s RAG handles longer documents. Neither replaces the discipline of knowing which lore entries a given scene actually needs.
- Continuity across sessions. Every session is a fresh context window unless the writer manages the summary. The trade-off is the same one novelists working without AI make: write the next scene, and keep notes on the one before. The model can help with the summary, but the writer owns the continuity.
A short decision framework
Five questions, in order:
- Are you OK with your manuscript being processed by a vendor LLM? If yes, NovelAI or Novelcrafter is faster to start. If no, the local stack is the only answer.
- How much VRAM do you have? 8 GB gets you Qwen3 4B and decent dialogue. 16 GB gets you Mistral Small 3.2 24B quantised or Qwen3 14B and confident prose. 24 GB unlocks Qwen3 30B-A3B or a full Mistral Small 3.2 24B. 64 GB+ gets you Llama 3.3 70B or a Qwen3 MoE quant.
- Do you need a character-card / lorebook UI? SillyTavern is the standard. If you want a notebook-style free-form generator instead, text-generation-webui.
- Do you also want to run a RAG pipeline over research notes? Open WebUI pairs the writing chat with a local knowledge base over uploaded documents.
- Do you want a single-file portable app? KoboldCpp is the only one of the four that ships as a standalone executable; the others want Node.js and a Python environment.
For most novelists on a laptop with 16 GB of RAM, the practical setup is Ollama running Mistral Small 3.2 24B (or Qwen3 14B), front-ended by SillyTavern for character work or Open WebUI for prose-and-notes. A workstation with 64 GB and a single 24 GB GPU can run Qwen3 30B-A3B for a quality bump. The bytes of the manuscript stay on the machine, the licence is permissive enough to publish with, and the front end is the one a writer actually wants to run on.