Best Local AI for Email: Triage, Draft, Reply Privately

Run an AI assistant against your inbox without sending mail to Google, Microsoft, or any third party. RAG setup and tradeoffs.

Your inbox is the most sensitive document you own. Bank alerts, medical confirmations, lawyer threads, HR complaints, source pitches, divorce paperwork - everything you would never print and hand to a stranger ends up there. The vendors that built email AI know this, and they have answered it differently: Microsoft says it does not train on your mail, Google lets you toggle smart features, with a default that depends on region, and the indie email apps range from “opt out in settings” to “we have not built the feature.” None of those answers involve the only entity that has no business interest in your correspondence: a model that runs on your own hardware.

This page walks through what “local AI for email” actually looks like in October 2026, what the tradeoffs are against hosted options like Microsoft 365 Copilot for Outlook, Gmail smart features, and Superhuman AI, and which combinations of local model and self-hosted front end cover the three jobs most people want - triage, drafting, and replying.

TL;DR

  • Local AI for email is two pieces, not one product. A self-hosted runner (Ollama, llama.cpp, LM Studio) supplies the model. A self-hosted interface (Open WebUI, AnythingLLM) supplies the inbox integration, RAG, and chat history. You can mix and match.
  • AnythingLLM ships dedicated Gmail and Outlook agents that talk to your mailbox via a user-owned Google App Script or Microsoft Graph app, not via a vendor relay. The LLM endpoint is whatever you configure - local or cloud.
  • The big hosted options have published privacy commitments but no opt-out from processing. Microsoft says “[p]rompts, responses, and data accessed through Microsoft Graph aren’t used to train foundation LLMs, including those used by Microsoft Copilot.” Google lets admins and individuals toggle smart features that include Gemini; the toggle is on by default outside the EEA, UK, Switzerland, and Japan.
  • If you want zero bytes leaving your machine, the path is Ollama + AnythingLLM (Gmail Agent or Outlook Agent) on the same host that holds your Ollama daemon. The full setup is described below.

What “local AI for email” actually means

Three tasks cover most email AI use:

  1. Triage. Summarise what is in the inbox, group threads, flag the items that need a human reply, archive the rest. This is read-only against the mailbox.
  2. Drafting. Produce a reply in your voice given a thread and a short instruction. This reads the thread and writes a draft to your drafts folder.
  3. Replying. Either send the draft automatically or hold it for your review. This is the highest-trust task: a runaway agent can ship email in your name.

A “local” deployment has to handle all three without sending the mailbox contents to a third-party LLM. The LLM itself has to run on hardware you control. The mailbox integration (the IMAP, Graph, or Gmail API calls) has to be authenticated as you, not as a vendor. The retrieval-augmented memory that lets the model answer “what did Alice say about the contract” has to live on the same hardware. Anything that bounces through someone else’s data centre breaks the chain, even if the model is open-weight.

Two pieces have to come together for this to work:

  • A model runner that keeps weights and inference on-device. Ollama is the most common pick; llama.cpp, LM Studio, and vLLM all qualify. Ollama’s own homepage states the position plainly: “Nothing you run locally ever leaves your machine.”
  • A front end that knows how to talk to a mailbox. This is the harder half. Two open-source projects have built real email integrations, and they take different approaches.

AnythingLLM is the most complete. Per the AnythingLLM docs, its Gmail Agent “can search emails, read messages and threads, compose and send emails, manage drafts, and organize your inbox,” and its Outlook Agent offers the same actions through a Microsoft Entra app. The Gmail Agent authenticates via a Google App Script that you copy into your own Apps Script editor, deploy as a Web App, and point at AnythingLLM with an API_KEY of your choosing. AnythingLLM never sees your Google password; it sees your Web App URL. Read actions run unattended; write actions marked with a pencil icon require your approval. The skill is single-user.

Open WebUI does not ship an email agent of its own in the same sense. Its local RAG supports nine vector databases and content-extraction engines, and it integrates with Google Drive and OneDrive/SharePoint for cloud document import. You can run an export of a mailbox into a folder, point Open WebUI at it, and ask the model questions about your own archive - but it is not a live mailbox integration. It is best understood as a “ask questions of email I already exported” tool, not a “draft a reply to the thread I have open” tool.

Both AnythingLLM and Open WebUI accept any OpenAI-compatible endpoint, so the same AnythingLLM Gmail Agent can run against a local Ollama model for full local operation, against a hosted model for a hybrid setup, or against a hosted model with a local Ollama fallback. The local-and-cloud flexibility is not a privacy regression - it is a knob.

What the hosted options actually do

A useful comparison needs the four products most readers will be choosing between, and what they actually say they do with mail.

Microsoft 365 Copilot for Outlook. Microsoft’s privacy documentation is unusually explicit. The Microsoft Learn page states: “Prompts, responses, and data accessed through Microsoft Graph aren’t used to train foundation LLMs, including those used by Microsoft Copilot.” Microsoft Copilot stores prompts and responses in the user’s Copilot activity history, which is encrypted at rest, governed by the same Microsoft 365 retention policy as the rest of the tenant, and can be deleted by the user via the My Account portal. For EU users the data stays inside the EU Data Boundary, with the caveat that “Models provided by Anthropic as a subprocessor are currently excluded from the EU Data Boundary.” Microsoft has also opted out of human review of Copilot content even though it is available in Azure OpenAI. The catch is that the prompts and responses are stored in your Microsoft 365 tenant (which an admin can read with Content search or Purview), and the model itself is hosted by Microsoft.

Gmail smart features and Gemini in Workspace. Google splits smart-feature controls into three toggles: Gmail/Chat/Meet, Workspace-wide (which covers Calendar, Voice, and Drive), and other Google products. The smart-features toggle is on by default everywhere except the EEA, the UK, Switzerland, and Japan, where it is off by default. Per Google’s help page on smart features, turning off a smart feature setting means “your Workspace Content & Activity will no longer be used to improve the relevant smart features moving forward.” Google’s wording is careful: “learnings developed from this improvement process may persist” even after the toggle is off. The opt-out is real, but the only way to get out of Google’s processing entirely is to keep smart features off across every product. In the EEA, UK, Switzerland, and Japan, smart features are off by default, which is a meaningful baseline. Outside those regions the default is opt-in to Gemini processing.

Superhuman. Superhuman’s privacy policy describes its data practices. User content - “emails and drafts, text, screen content, web pages, documents, files, calendars, images, data” - is collected, and the policy says “you can decide whether Superhuman can use your user content to train our AI models by adjusting the available training control(s) in your account settings.” The policy does not state the default setting of that control. The policy is also explicit on what is not done with user content: “We do not use your user content for marketing or advertising purposes.” Superhuman is a closed-source hosted product; there is no local option. The catch is the same as Gmail’s: an opt-out exists, but the only way to ensure nothing leaves is to not use the product.

Shortwave and other indie AI email apps. Shortwave, SaneBox, Mailmeteor, and the rest of the long tail of AI-forward email clients each have their own privacy posture, and the cluster is moving quickly enough that pointing at one is risky. If the question is “what does the AI company see,” the answer is “the prompt, the message it is replying to, and the model’s output” unless the privacy policy says otherwise.

The table below puts the four side by side on the questions a privacy-first reader cares about. The “training” column is “trains by default on user content,” and the “user data leaves your machine” column is “bytes go to a vendor under normal operation.”

ToolTrains by default on mailUser data leaves your machineLocal-only modeOpen-source
Microsoft 365 Copilot for OutlookNo (foundation LLMs not trained on prompts, responses, or Graph data)Yes (Microsoft cloud)NoNo
Gmail smart features / GeminiYes when smart features are on (off by default in EEA, UK, Switzerland, Japan)Yes (Google cloud)NoNo
Superhuman AIUser content can be used for training; a control in account settings governs it (default not stated in the policy)Yes (Superhuman cloud)NoNo
AnythingLLM + Ollama (Gmail / Outlook Agent)No (model runs on your host; no vendor sees mail content)No (mailbox authentication is user-owned App Script or Graph app)YesYes (AnythingLLM is MIT; Ollama is MIT)

The table is the core claim of this page: every hosted option processes mail on vendor servers, and where training is possible the user has to act to control it. The local option keeps model inference on your own hardware, so no LLM vendor sees the prompts or the drafts.

Setting up local AI for email: the minimum viable stack

A working local-email-AI setup, end to end:

  1. Pick a model. For triage and short replies, Ollama’s qwen3:4b (2.5 GB / about 2.3 GiB download) or mistral-small:24b (14 GB / about 13.0 GiB download) handles email-thread-length context and writes in plain English. For longer threads or more careful drafts, qwen3:30b-a3b (19 GB / about 17.7 GiB download) or llama3.3:70b (43 GB / about 40.0 GiB download) is the next step up. Download size is a floor, not the full memory need: context adds to it.
  2. Run Ollama locally. ollama serve listens on 127.0.0.1:11434 by default. No network calls leave the host.
  3. Install AnythingLLM. AnythingLLM ships a desktop installer and a Docker image. Point AnythingLLM’s LLM setting at the Ollama endpoint and pick the model. AnythingLLM is MIT-licensed.
  4. Wire up the Gmail Agent or Outlook Agent. The Gmail Agent setup page walks through copying a Google App Script into your Apps Script editor, deploying as a Web App “Execute as: Me, Access: Anyone,” pasting the deployment ID and an API_KEY you choose into AnythingLLM. The Outlook Agent uses a Microsoft Entra app registration with Graph permissions.
  5. Use it. From AnythingLLM’s chat, you can now ask “summarise the last 24 hours of unread inbox,” “draft a reply to Alice’s contract thread saying we accept the indemnity clause,” or “list everything from Bob that I have not replied to in a week.” AnythingLLM sends the action to the App Script, the App Script calls Gmail as you, the model is your local Ollama model, and no LLM vendor sees the thread (Google or Microsoft still hosts the mailbox).

The setup takes about an hour if you have used Google App Script or Microsoft Graph before, longer the first time. The privacy delta is large: every byte that would have gone to a vendor LLM now stays on your hardware.

What the local stack does not solve

Three honest limitations.

  1. Quality on small models. A 4B model will draft a polite reply that says “I will follow up next week” when you wanted “I will not be able to attend.” Larger local models close the gap but the gap exists. Picking the right model for the right task is the work of the local-AI cluster on this site.
  2. Live mailbox freshness. The Gmail and Outlook agents operate on the live mailbox through the API; they do not maintain a private index. AnythingLLM’s workspace memory and RAG layer can hold prior threads if you import them, but a model that says “what did Alice say about the contract last month” depends on that import having run. Live access is via the agent.
  3. Approval loops. The Gmail Agent marks destructive operations with a pencil icon that requires your explicit approval, which is the right default. It does not solve “I forgot to approve and the agent sent something I would not have sent.” If the workflow is “draft only, never send,” the right AnythingLLM setting is to disable the send actions entirely.

A short decision framework

Five questions, in order:

  1. Are you OK with mail being processed by a vendor LLM? If yes, hosted is faster to start. If no, the local stack is the only answer.
  2. Do you trust Microsoft / Google / Superhuman’s opt-out controls? If yes, hosted with the opt-out enabled is fine. If no, again, local.
  3. Are you inside the EEA, UK, Switzerland, or Japan? Google’s smart features are off by default there; Microsoft’s commitments are global. Outside those regions, Google’s default is opt-in to Gemini processing.
  4. Are you drafting replies in a sensitive domain - law, medicine, finance, journalism? Hosted models have a non-zero history of prompt-injection incidents, and a vendor LLM in your inbox is a high-value target. Local closes the attack surface: nothing leaves the machine.
  5. Are you OK with the quality delta of a 4B to 24B local model? If yes, local from the start. If no, hybrid: AnythingLLM with Ollama as the default, falling back to a hosted model for the threads where quality matters, with the fall-back traffic going to a model whose privacy posture you accept.

For most readers with a laptop that has 16 GB of RAM and a 4B to 14B Ollama model, the hybrid setup covers triage and most drafts. A local 70B quant on a workstation covers everything. The only thing hosted covers better is the very long horizon, very subtle context that needs a frontier model - and that is the only thing you cannot get locally in 2026 without a multi-GPU rig.