A cover letter is a paragraph about a person, written to a specific role at a specific company, with the job ad pasted next to it. A resume is a structured list of the same person’s work. Both are loaded with personal data: name, address, past employer names, manager names, dates, salary hints, gaps in employment, schools, professional contacts, and sometimes medical or family context in a “cover note”. When that text gets sent to a hosted AI assistant, it is processed under that vendor’s retention and training policy, which is rarely what a job applicant wants. Local models keep the document on the machine: the prompt, the output, and the iteration loop never leave the laptop. This guide walks through which open-weight model to pick by RAM, what each model does well on prose drafting, and how to wire it up so the cover letter workflow runs offline.
TL;DR
- The three properties that matter for cover letter work are: instruction following on long generation, role-specific tone control, and licence. A cover letter is usually 250 to 400 words; a resume bullet is one to two lines; both need the model to follow a format you set.
- For 8 GB to 16 GB laptops without a GPU,
qwen3:4bat 2.5 GB (Apache-2.0) is the practical default. It generates clean prose, handles a system prompt that sets a tone (“write as a senior backend engineer”), and stays under 16 GB of RAM with room to spare. - For 16 GB to 32 GB laptops with an entry GPU,
qwen3:8bat 5.2 GB is the realistic pick. The Qwen3-8B card lists 8.2B total / 6.95B non-embedding parameters, 32,768 tokens native context (131,072 with YaRN), and “strong alignment for creative writing, role-playing, multi-turn dialogues, and instruction following” at Apache-2.0. - For a 24 GB GPU or a 32 GB unified-memory Mac,
mistral-small:24bat 14 GB (Apache-2.0 per the Mistral-Small-24B-Instruct-2501 card) gives better long-prose quality. For the smallest practical footprint,gemma3:4bat 3.3 GB with a 128K context (Gemma Terms of Use) is the alternative for lighter hardware, with Google’s Gemma card listing 140+ languages. - Use a chat interface or the Ollama CLI to keep the loop private. The Ollama API docs describe the
systemparameter on/api/generate(which “overrides what is defined in the Modelfile”) and the system message role on/api/chat; both keep prompts off vendor servers when the model is local. - Licence matters if you ship a product, less so for personal job-search drafts. Apache-2.0 (Qwen3, Qwen2.5, Mistral Small) and MIT (Phi-4-mini-instruct) are permissive. Gemma 3 ships under Google’s separate Gemma Terms of Use, which permits commercial use but requires the use-restriction notice to flow downstream. Llama 3.2 3B ships under Meta’s community licence.
What “cover letter and resume drafting” actually requires
Three capabilities separate a model that can write a coherent cover letter from one that returns generic prose.
- Long-form instruction following. A cover letter prompt typically includes a job ad, a short background summary, and instructions about tone and length (250 to 400 words, three paragraphs, avoid clichés). The Qwen2.5-7B-Instruct card lists “significantly improved instruction following, generating long texts (over 8K tokens), and more resilient to the diversity of system prompts” as headline Qwen2.5 gains, and “generating structured outputs especially JSON” alongside it. The Qwen3-8B card keeps those strengths and adds a hybrid thinking/non-thinking mode via
enable_thinking. - Role-specific tone control. A cover letter for a senior backend role reads differently from one for an entry-level designer. The system prompt is where the tone lives, and the model has to follow it. The Ollama API docs describe two paths: pass
systemas a field on/api/generate, or include a{"role": "system", "content": "..."}message in themessagesarray on/api/chat. Both let the user set the voice before any document content goes in. - Structured output for resume work. Resume bullets fit a fixed shape (action verb, task, outcome, metric). The Ollama API
formatparameter accepts"json"for JSON mode or a JSON Schema for structured outputs that constrain the model to a shape, on both/api/generateand/api/chat. A schema is what makes resume output land in the same column every time.
A fourth, less obvious capability matters: the ability to take the user’s existing CV or cover letter as input and rewrite it without losing the original voice. Long-context models handle this on one prompt; shorter-context models need chunking. For 250-word cover letters, context length is rarely the binding constraint; for resume rewrites with five to ten years of work history, it starts to be.
The 2026 open-weight field for cover letter and resume work
Five families cover the realistic picks for prose drafting on a laptop or a small server. They differ more on instruction tuning and licence than on raw capability at the same parameter count.
Qwen3 (Apache-2.0, Alibaba)
The Ollama qwen3 library page ships tags from 0.6B to 235B. The relevant prose-drafting tags fit a wide range of hardware:
| Tag | Download | Context |
|---|---|---|
qwen3:1.7b | 1.4 GB | 40K |
qwen3:4b | 2.5 GB | 256K |
qwen3:8b | 5.2 GB | 40K |
qwen3:14b | 9.3 GB | 40K |
All are Apache-2.0. The Qwen3-8B card lists 8.2B total / 6.95B non-embedding parameters, 32,768 tokens native context (131,072 with YaRN), and explicit strengths in “creative writing, role-playing, multi-turn dialogues, and instruction following.” The Qwen2.5-7B-Instruct card is the prior-generation model with the same instruction-following focus and a 131,072-token context window. For cover letter work the practical picks are qwen3:4b for the laptop default and qwen3:8b for the entry GPU; qwen3:1.7b covers the floor.
Qwen2.5-7B-Instruct (Apache-2.0, Alibaba)
The Qwen2.5-7B-Instruct card describes a 7.61B-parameter causal LM (6.53B non-embedding, 28 layers, 28 Q / 4 KV attention heads, GQA) with a 131,072-token context and “significantly improved instruction following, generating long texts (over 8K tokens).” It is the right pick when the cover letter work needs longer generation runs than Qwen3 8B’s default and the user has a 16 GB-class machine.
Phi-4-mini-instruct (MIT, Microsoft Research)
The Phi-4-mini-instruct card describes a 3.8B-parameter dense decoder-only Transformer with 128K context, MIT licensed, designed for “memory/compute-constrained environments and latency-bound scenarios.” The card lists an overall aggregate benchmark score of 63.5 against Qwen2.5-3B-Ins at 60.1, Mistral-8B-2410 at 60.2, and Llama-3.1-8B-Ins at 62.3. The card documents a function-calling format with <|tool|> tokens; for cover letter work, Phi-4-mini is the alternative when the user wants a smaller model with explicit chat-format tooling.
Gemma 3 (Gemma Terms of Use, Google DeepMind)
The Google gemma-3-4b-it card describes a 4B-parameter multimodal vision-language model with 128K input / 8K output context, 140+ language support, and trained on 4 trillion tokens. The licence is Google’s separate Gemma Terms of Use, which permits commercial use, modification, and distribution but requires the use-restriction notice from Section 3.2 (incorporating the Gemma Prohibited Use Policy) to flow to downstream recipients. For personal cover letter work the licence is rarely a blocker; for any product built on Gemma, the redistribution terms matter.
Mistral Small 24B (Apache-2.0, Mistral AI)
mistral-small:24b is 14 GB on Ollama with a 32K context window, Apache-2.0 per the Mistral-Small-24B-Instruct-2501 card. At 24B parameters it produces noticeably better long-prose quality than the 4B to 8B tier, and is the right pick on a 24 GB GPU or a 32 GB unified-memory Mac. The 32K context is enough for a job ad plus a multi-paragraph cover letter on one prompt.
Llama 3.2 3B (Llama 3.2 Community License, Meta)
llama3.2:3b is 2.0 GB on Ollama at Q4_K_M, 3.21B parameters, 128K context, Llama 3.2 Community License. It is the smallest mainstream pick and runs on a Raspberry Pi or a 4 GB RAM machine. Its prose is serviceable but not as instruction-tuned as Qwen3 or Qwen2.5; the licence has a separate commercial-use clause worth reading before relying on it for anything beyond personal drafts.
Pick by hardware tier
4 GB to 8 GB RAM, no GPU
qwen3:1.7b (1.4 GB, Apache-2.0) covers one-shot drafts: a 250-word cover letter with a short system prompt fits in well under 2 GB of weights and KV cache. Quality is serviceable for short drafts; the model tends to drift on long generation.
8 GB to 16 GB RAM, no GPU or small integrated GPU
qwen3:4b at 2.5 GB (Apache-2.0) is the practical default. It fits comfortably in RAM on an 8 GB machine and leaves room for the Ollama server itself. The Qwen3-8B card lists the Qwen3 family’s creative-writing and instruction-following gains; the 4B reference flows from the same family. For longer generation runs, gemma3:4b at 3.3 GB (128K context, Gemma Terms of Use) is the alternative.
16 GB to 32 GB RAM with an 8 GB to 12 GB GPU (RTX 3060, Apple M-series shared memory)
qwen3:8b at 5.2 GB (Apache-2.0 per the Qwen3-8B card) is the realistic pick. The card lists 32,768 tokens natively (131,072 with YaRN), “instruction following, mathematics, code generation, and tool calling” strengths, and Apache-2.0. Qwen2.5-7B-Instruct at 7.61B parameters (131K context, Apache-2.0) is the alternative when the user needs longer generation runs than Qwen3 8B’s default.
24 GB GPU (RTX 4090) or 32 GB unified memory (Mac Studio)
mistral-small:24b at 14 GB (Apache-2.0 per the Mistral-Small-24B-Instruct-2501 card) is the strongest small-server pick. Long-prose quality is noticeably better than the 8B class, and 32K of context fits a job ad plus a multi-paragraph cover letter on one prompt.
A practical local setup
Three steps get a local cover letter workflow running without any cloud round-trip:
- Install Ollama and pull a model.
ollama pull qwen3:8bdownloads 5.2 GB to~/.ollama. The Ollama API docs describe the server’s default endpoint atlocalhost:11434. - Use a chat interface or the CLI.
ollama run qwen3:8bopens a local REPL. For a multi-document workflow, Open WebUI on the same machine gives a sandboxed chat client that talks to the local Ollama server without any external routing. Either way, the prompt and the output stay on the host. - Set the system prompt per role. The Ollama API docs describe the
systemparameter on/api/generate(which “overrides what is defined in the Modelfile”) and the system message role on/api/chat. A reusable prompt per role (“write as a senior backend engineer applying to a Series B startup; avoid AI cliches; close with a single paragraph”) turns the workflow into a template.
For structured resume work, pass "format": "json" for JSON mode or a JSON Schema to the format parameter. The Ollama API describes how the server constrains the model’s output to the schema on both endpoints. The Qwen2.5-7B-Instruct card lists “generating structured outputs especially JSON” as a headline Qwen2.5 improvement, which the Qwen3 family inherits.
What to skip
Four things do not work well for cover letter and resume work and are worth calling out.
- Tiny models without instruction tuning. Anything below 1.5B parameters tends to drift off-tone and ignore the format you set.
qwen3:1.7bis the floor for serious prose work; below that, treat the output as a draft to rewrite, not a draft to send. - Hosted AI for confidential applications. The point of running locally is that the cover letter never leaves the machine. A hosted ChatGPT or Claude session processes the prompt under the vendor’s retention policy, and the resume text often flows into a training pipeline unless the user has explicitly opted out in that product. Local does not solve every problem, but it solves this one cleanly.
- Long-context models when the workload is short. A 128K-context Gemma 3 or Phi-4-mini is overkill for a 250-word cover letter. The KV cache cost is wasted on work that fits in 4K. Pick by RAM tier, not by context window alone.
- Trusting tone without reading. A local model can still write “I am uniquely qualified” the same way a hosted model can. Read the output before you send it. The model is a drafting tool, not a sender.
The bottom line
For cover letter and resume work on an 8 GB to 16 GB laptop, qwen3:4b at 2.5 GB (Apache-2.0) is the default, with a system prompt that sets the tone. For 16 GB RAM plus an entry GPU, qwen3:8b at 5.2 GB (Qwen3-8B card) is the realistic pick. For a 24 GB GPU or a 32 GB unified-memory Mac, mistral-small:24b at 14 GB (Mistral-Small-24B-Instruct-2501 card) gives the strongest long-prose quality. gemma3:4b at 3.3 GB (Gemma Terms of Use) is the alternative when 140+ language support matters more than raw quality. The Ollama API system parameter and the format parameter for JSON output are what keep the workflow private and the resume bullets in the same column. Running locally matters if the cover letter cannot leave the building.
Related on Intelligibberish
- How to Redact Personal Info Before Pasting Into an AI Assistant - the privacy pre-flight for when local is not an option.
- Best Local LLM Runners: Ollama, LM Studio, llama.cpp, vLLM - the four main ways to serve the weights you pick here.
- Tools - more AI-tools reviews, comparisons, and hands-on guides.