Best Local LLM for Spreadsheet Analysis on a Laptop

Practical tier guide to running a local LLM on Excel, CSV, and tabular data. Picks by RAM, JSON output for clean rows, and what to skip.

Spreadsheet work is the kind of task where uploading the file is the privacy decision: a CSV of payroll, a workbook of customer accounts, a quarterly close in Excel. Each of those rows has at least one field that should not be sent to a vendor. Local models let you ask “group these transactions by counterparty and flag anything over $10,000”, or “summarise this month’s P&L by category”, without the file leaving the machine. This guide walks through which open-weight model to pick by RAM, how to force the model to return JSON you can paste back into a sheet, and the trade-offs between reasoning depth and wall-clock speed.

TL;DR

  • The three properties that matter for spreadsheet work are: structured-output fidelity (the model returns valid JSON that parses), instruction following on tables, and licence. Then check how much RAM you have.
  • For JSON-output work the Qwen2.5 generation is the practical default. The Qwen2.5-7B-Instruct card states that Qwen2.5 brought “significant improvements in” instruction following, “generating structured outputs especially JSON”. The Qwen2.5-Coder-7B-Instruct card keeps that focus and adds 131K context for long sheets.
  • For small laptops (8 GB to 16 GB RAM, no GPU), qwen3:4b at 2.5 GB (Apache-2.0) is the default. qwen3:1.7b at 1.4 GB covers light “summarise this column” work and runs on 8 GB machines comfortably.
  • For 16 GB to 32 GB laptops with a small GPU, qwen3:8b at 5.2 GB is the realistic pick. The Qwen3-8B card lists 8.2B parameters, 32,768 tokens native context (131,072 with YaRN), and Apache-2.0. Phi-4-mini-instruct at 3.8B (MIT, 128K context) is the alternative if you want a smaller-footprint model with explicit JSON tool-call support.
  • For a 24 GB GPU or a 32 GB unified-memory machine, qwen3:30b-a3b at 19 GB is the strongest small-server pick. The Qwen3-30B-A3B-Instruct-2507 card lists 30.5B total / 3.3B active parameters, 262,144-token native context, and Apache-2.0.
  • Use the Ollama format parameter to force JSON. The Ollama API docs describe JSON mode (format: "json") and structured outputs with a JSON schema; both work on /api/generate and /api/chat. A schema is what makes output pasteable into a sheet column.
  • Licence matters for commercial work. Apache-2.0 (Qwen3, Qwen2.5, Qwen2.5-Coder) and MIT (Phi-4-mini) are permissive; Llama 3.2 (Ollama tag page) ships under Meta’s community licence.

What “spreadsheet analysis” actually requires

Four capabilities separate a model that can answer a question about a sheet from one that returns prose.

  1. Structured-output fidelity. The single most common failure mode is a model that returns “Here is the JSON: …” followed by a code block that is not JSON. The Qwen2.5 generation card calls out the family-wide improvement here: the Qwen2.5-7B-Instruct card lists JSON-structured-output generation among the headline gains of the Qwen2.5 family. The Ollama API lets you go further: pass a JSON schema in the format parameter and the server constrains the response to match it, on both /api/generate and /api/chat.
  2. Tool-call format. Some spreadsheet workflows route the model through a function-call interface (“call sum_column(sheet=Sheet1, column=Amount)”) rather than parsing free-form text. The Phi-4-mini-instruct card describes a function-calling format: tools are specified as JSON inside <|tool|>…<|/tool|> tokens in the system prompt, and the model emits structured function calls.
  3. Long-context behaviour on rows. A 50,000-row CSV is short in tokens but long in context for an LLM: each cell is a few tokens, so the whole sheet can be tens of thousands of tokens. The Qwen3-8B card lists 32,768 tokens natively and 131,072 with YaRN; Phi-4-mini lists 128K native. A model with RULER-style long-context evaluation is more likely to find a row at the bottom of a wide sheet when you ask about it.
  4. Licence. Apache-2.0 (Qwen3-8B, Qwen2.5-7B-Instruct, Qwen3-30B-A3B-Instruct-2507) and MIT (Phi-4-mini) are permissive. Llama 3.2 3B ships under the Llama 3.2 Community License, which has commercial-use terms you should read before shipping a product.

The 2026 open-weight field for spreadsheet work

The relevant picks split into four families. They differ more on instruction tuning than on raw capability at the same parameter count.

Qwen3 (Apache-2.0, Alibaba)

The Ollama qwen3 library ships tags from 1.7B to 235B. The smaller tags fit a wide range of hardware:

TagDownloadNative context
qwen3:1.7b1.4 GB32K
qwen3:4b2.5 GB256K
qwen3:8b5.2 GB32K
qwen3:14b9.3 GB32K
qwen3:30b-a3b19 GB256K

All five are Apache-2.0. The Qwen3-8B card lists 8.2B total / 6.95B non-embedding parameters, a 32,768-token native context extendable to 131,072 with YaRN, hybrid thinking/non-thinking with enable_thinking, support for 100+ languages, and explicit strengths in instruction following, math, code generation, and tool calling. The Qwen3-30B-A3B-Instruct-2507 card lists 30.5B total / 3.3B active parameters, 262,144-token native context, and “substantial gains in long-tail knowledge across multiple languages”. For spreadsheet work the practical picks are qwen3:4b, qwen3:8b, and qwen3:30b-a3b for the GPU tiers; qwen3:1.7b is the floor.

Qwen2.5-Coder (Apache-2.0, Alibaba)

The Coder generation is tuned for code and structured output. The Qwen2.5-Coder-7B-Instruct card describes a 7.61B-parameter causal LM focused on “code generation,” “code reasoning,” and “code fixing,” with a 131,072-token context. The Qwen2.5-Coder-14B-Instruct card is the next step up: 14.7B parameters, 48 layers, 40 query heads / 8 KV heads (GQA), 131,072-token context. The Coder generation is the right pick when the spreadsheet task looks like code: generating a formula, writing a Python pandas snippet to reshape a CSV, or producing a SQL query against a database.

Phi-4-mini-instruct (MIT, Microsoft Research)

The Phi-4-mini-instruct card describes a 3.8B-parameter dense decoder-only Transformer with 128K context, MIT licensed, designed for “memory/compute-constrained environments and latency-bound scenarios.” The card reports GSM8K 88.6, MATH 64.0, and an Overall score of 63.5 against Llama-3.1-8B-Ins at 62.3. For structured output the card documents a function-calling format: tools specified as JSON inside <|tool|>…<|/tool|> tokens in the system prompt. If you are running a tiny local model and want explicit JSON tool calls rather than free-form output, Phi-4-mini is the more conservative pick than Qwen2.5-Coder.

Llama 3.2 (Llama 3.2 Community License, Meta)

llama3.2:3b is 2.0 GB on Ollama at Q4_K_M, 3.21B parameters, Llama 3.2 Community License. It is the smallest mainstream pick and runs comfortably on a Raspberry Pi or a 4 GB RAM machine. It is not tuned for structured output the way Qwen2.5 is, and the licence’s commercial-use terms deserve a careful read before you ship a product built on it.

Pick by hardware tier

4 GB to 8 GB RAM, no GPU

qwen3:1.7b (1.4 GB, Apache-2.0) covers one-shot tasks on small sheets: “summarise this column”, “translate these product names”, “categorise these expenses”. For anything beyond simple prompts, move up to qwen3:4b.

8 GB to 16 GB RAM, no GPU or small integrated GPU

qwen3:4b at 2.5 GB (Apache-2.0) is the practical default. It fits comfortably in RAM on an 8 GB machine and leaves room for the Ollama server itself. For JSON output, pass a JSON schema in the Ollama format parameter. Phi-4-mini-instruct at 3.8B (MIT) is the alternative if you specifically want function-call output.

16 GB RAM plus an 8 GB to 12 GB GPU (RTX 3060, Apple M-series shared memory)

qwen3:8b at 5.2 GB (Apache-2.0, Qwen3-8B card lists 8.2B parameters and 32K native / 131K YaRN context) is the realistic pick for a laptop with an entry GPU or enough unified memory. Qwen2.5-Coder-14B-Instruct at 14.7B parameters is the alternative for code-shaped work (formula generation, pandas snippets) if you can fit its weights and KV cache.

32 GB RAM with a 24 GB GPU (RTX 4090) or 64 GB+ unified memory (Mac Studio)

qwen3:30b-a3b at 19 GB (Apache-2.0, Qwen3-30B-A3B-Instruct-2507 card lists 30.5B total / 3.3B active and 262,144-token native context) is the strongest small-server pick. MoE with 3.3B active means the speed is closer to a 4B model than a 30B one, while the long context window lets you fit large CSVs or several sheets at once.

How to force the model to return clean JSON

The single biggest productivity gain on spreadsheet work is teaching the model to return structured output you can paste into a column. The Ollama API docs describe two ways:

  1. JSON mode. Pass "format": "json" in the request body. The model returns a valid JSON object. Use this when the schema is implicit in the prompt (for example, “respond with JSON, keys: question, confidence”).
  2. Structured outputs with a JSON schema. Pass a JSON Schema as the format value. The server constrains the model’s output to match the schema. Use this when every row must have the same shape (a “category” column, an “amount” column, a “flag” column). The docs describe this on both /api/generate and /api/chat.

The Qwen2.5 generation is the family that most reliably respects the format parameter. The Qwen2.5-7B-Instruct card lists “generating structured outputs especially JSON” as a headline Qwen2.5 improvement, and the Qwen2.5-Coder generation carries the same focus. If you find the model wrapping JSON in markdown code blocks anyway, set the system prompt to “respond with raw JSON, no markdown, no commentary” and validate with jq on the way out.

For function-call workflows, the Phi-4-mini-instruct card describes the alternative: tools are specified as JSON inside <|tool|>…<|/tool|> tokens in the system prompt, and the model emits structured function calls you can dispatch directly.

What to skip

Three patterns do not work well for spreadsheet work and are worth calling out.

  1. Long-form prose answers. A model that returns “Here is a summary of your sheet: …” is unusable in a column. Force JSON mode or write a tool-call interface; do not paste the model’s prose back into a sheet.
  2. Tiny models without JSON tuning. Anything below 1.5B parameters tends to fail the JSON-schema constraint under load. qwen3:1.7b is the floor for serious JSON-mode work; below that, treat the output as a draft to validate, not as final.
  3. Hosted AI for confidential sheets. The point of running locally is that the file never leaves the machine. A hosted ChatGPT or Claude session processes the sheet under the vendor’s retention policy, and a long spreadsheet often triggers the vendor’s training pipeline unless the user has explicitly opted out in that product. Local does not solve every problem, but it solves this one cleanly.

The bottom line

For spreadsheet analysis on an 8 GB to 16 GB laptop without a GPU, qwen3:4b (2.5 GB, Apache-2.0) is the default, and the Ollama API format parameter is what makes the output pasteable. For 16 GB RAM plus an entry GPU, qwen3:8b is the realistic pick. For a 24 GB GPU or a Mac Studio with unified memory, qwen3:30b-a3b at 19 GB gives the strongest instruction following and 256K context. Phi-4-mini-instruct is the alternative when function-call output matters more than raw quality. The licence (Apache-2.0 or MIT) matters if you build on the model; running locally matters if the sheet cannot leave the building.