AI Chats That Accept Image Uploads Without Training on Them

Which AI chats accept image uploads without training on your photos, what they retain, and the local fallback.

Photos carry information text never does. A photo of a tax document shows your employer, your income, and your face. A photo of a screenshot leaks the URL bar and the open tab. A photo of a whiteboard carries handwriting that OCR turns into searchable text. A photo of your kid carries their face, their school uniform, and the location metadata embedded by your phone.

That is why the question “can I upload an image to an AI chat” is its own privacy question, separate from “can I send text”. Every vision-capable chat handles the upload, the temporary processing, the retention, and the training pipeline differently, and almost none of them publish a setting that is specifically about images.

This page is a privacy-first comparison of the hosted AI chats that actually accept image uploads today, plus the local fallback that beats them on all of the above.

TL;DR

  • Image uploads are a different threat surface from text. Photos carry EXIF metadata (GPS, device, timestamps), embedded text that OCR can lift, and the biometric identifiers in faces and handwriting. Choosing a chat on the basis of its text-chat policy alone misses the upload itself.
  • Anthropic and Google both train on images by default, and both offer an opt-out, but the opt-out has to be set in advance and does not retroactively strip data already seen. Anthropic’s opt-out path is documented; Google’s depends on the Keep Activity toggle.
  • Duck.ai and Venice strip personal metadata before the image leaves your device. Duck.ai binds every request to a contracted provider that is required to delete within 30 days and is prohibited from using prompts and text to develop models.
  • Ollama with a vision model and Open WebUI gives you a fully local path. The image never leaves your machine, the only retraining is whatever you do on your own weights, and the privacy work is the same as the local text chat you may already run.

What each hosted chat actually does with your uploaded image

This section compares what each vendor’s own privacy page says. Quotes are taken from the page cited in the source list. Where a vendor does not publish an image-specific statement, that absence is noted - it does not mean the image is handled better.

Anthropic Claude (claude.ai)

Claude accepts image, PDF, and other file uploads on free, Pro, Max, and Claude Code plans. The privacy policy treats any content you submit as an “Input,” and the training position is the same for images as for text. Anthropic writes that it “may use your Inputs and Outputs to train and improve Anthropic AI models, unless you opt out.”

The opt-out exists and is documented. Anthropic’s help center walks through the path: open claude.ai/settings/profile, then go to Privacy (claude.ai/settings/data-privacy-controls), and toggle off “Help Improve our AI models.” On mobile, the path is name -> Settings -> Privacy. The page is explicit about what the toggle does and what it does not. Anthropic says “we will not use any new chats and coding sessions you have with Claude for future model training,” and adds that chats already used in an active or completed training run stay in that run.

Two exceptions do not respond to the toggle. Anthropic notes that flagged conversations may still be used for trust-and-safety models, and that explicitly reported feedback is reviewed regardless of the setting. The retention language is general: Anthropic keeps data “for as long as reasonably necessary” and removes deleted conversations from history immediately and from back-end systems within 30 days.

Net position for images: opt out before you upload anything you would not want a future grader to read.

Google Gemini (gemini.google.com)

Gemini accepts image uploads on the web and Android, and on iOS via the Gemini app. The Gemini Apps Privacy Hub treats uploaded images as content for the prompt, and notes that Gemini uses Google Lens to read text inside the image. Google says Gemini Apps “use this information just like any other prompt.”

When Keep Activity is on, the page says uploaded images may be used to improve Google services with the help of human reviewers. Reviewed chats are disconnected from your account before reaching reviewers and are retained for up to 3 years. The default for Activity is auto-delete after 18 months; 3-month, 36-month, and indefinite windows are configurable. With Keep Activity off, future chats are saved for 72 hours. Temporary chats are not used to train AI models and are retained for 72 hours with your account.

Two practical consequences for images. First, “use this information just like any other prompt” means an OCR-readable screenshot of a sensitive email is treated no differently from a baseline chat. Second, when Keep Activity is on, the path from upload to human review is real and the disconnector is not a guarantee that the image is anonymous - it strips account link, not the content itself.

DuckDuckGo Duck.ai (duck.ai)

Duck.ai is the privacy-fronted chat wrapper. DuckDuckGo’s privacy page says it “calls model providers on your behalf” so the provider does not see your IP address. It says “All metadata that contains personal information (for example, your IP address) is removed before sending Prompts to underlying model providers,” and that providers are required to delete data “at most within 30 days, with limited exceptions for safety and legal compliance.” DuckDuckGo’s stated model providers include Anthropic, Azure OpenAI, OpenAI, and together.ai.

The catch is that DuckDuckGo is a wrapper, not a model, so the underlying provider’s training policy still applies for whatever the provider keeps. DuckDuckGo’s contracts say providers may not use prompts and text to develop or improve their models, and that the provider has to delete within 30 days. DuckDuckGo does not say in the page I read whether Duck.ai accepts image uploads in its current interface, so the image-specific position is worth verifying on the live UI before you assume it.

Venice AI (venice.ai)

Venice routes to a long list of vision-capable models. The Venice Models page lists vision support across many entries (Claude Sonnet, GPT, Gemini, Llama 4, Gemma 3, Qwen 3.5/3.6 Vision, Qwen3 VL 235B/30B, Mistral Small 3.2, and others), with a “Vision” tag in the Capabilities column. The privacy docs describe four privacy modes - Anonymous, Private, TEE, and E2EE - and add a critical caveat: “TEE and E2EE are currently available on text models only.” That is, the strongest Venice privacy guarantees do not apply to the image you upload.

Venice’s main privacy policy treats Prompts and Outputs as content Venice does not store. It says “We may record your actions (e.g., that you created a chat), but we will not have access to or store the content of your Prompts or Outputs,” and that Venice has a zero-data-retention arrangement with its model providers. The privacy policy’s section on likeness uploads (a separate feature) discusses BytePlus Singapore and a one-hour upload-storage window, but does not explicitly state whether chat-time vision image uploads are used to train Venice’s models. Read the page yourself before assuming the chat-image position is the same as the text position.

Mistral Le Chat and Proton Lumo

Both products list vision-capable models on their feature pages. The privacy documents I read in full did not contain a paragraph specific to image uploads in chat - they discuss retention and training at the chat level, not the upload level. Treat that as a known gap. For sensitive image uploads, the honest position is to assume the same retention and training rules as text until the vendor publishes an image-specific paragraph, and to verify by reading the vendor’s policy directly on the day you upload.

The local fallback: Ollama + a vision model + Open WebUI

A local vision model removes the upload itself from the threat model. The image leaves your machine only over the loopback interface to the local inference server, and no remote service ever receives the pixels.

Ollama’s API documentation supports image inputs. The docs show the images parameter for both /api/generate and /api/chat, accepting base64-encoded images for multimodal models. Ollama’s multimodal blog lists the models it treats as first-class vision targets: Meta Llama 4, Google Gemma 3, Qwen 2.5 VL, and Mistral Small 3.1. The blog is explicit that the multimodal engine was rebuilt because supporting vision models on the legacy llama.cpp path “became more and more challenging,” and now supports document scanning, OCR, and translation.

In practice, the recipe is: install Ollama, pull a vision-capable model from Ollama’s library (the multimodal blog lists Meta Llama 4, Google Gemma 3, Qwen 2.5 VL, and Mistral Small 3.1, and the API documentation uses the older LLaVA family as its example), install Open WebUI, and chat through the web UI. Open WebUI passes uploaded images to the local model via Ollama’s API. The image never leaves your machine.

Three honest limits on the local path. First, model quality on image reasoning still trails the largest hosted frontier vision models on hard tasks; OCR and casual image chat are well-served, exotic charts and dense documents less so. Second, you inherit Ollama’s security surface: the API binds to localhost by default, but if you bind it to a LAN address without authentication, anyone on your network can hit it. Third, the privacy work you do not skip is still yours: the image is on your disk, the EXIF metadata is still in the JPEG, and the screenshot is still in your clipboard. Local inference is not a substitute for redacting the photo before you decide to keep it.

A pre-flight checklist before any image upload

Use this before uploading any photo to any hosted AI chat.

  1. Strip EXIF before uploading. Photos carry embedded location, device, and timestamp metadata that travels with the file unless you explicitly remove it. Use your operating system’s photo viewer or a tool like exiftool to strip GPS and device fields before the upload. EXIF on its own can leak GPS coordinates down to a few meters.
  2. Decide whether the image has readable text. If it is a screenshot of a sensitive email, an internal dashboard, a tax form, a medical record, or a code editor, the text inside is just text. OCR on the receiving side treats it like a chat. Crop to the relevant area, or upload a screenshot with the URL bar and sidebars blurred.
  3. Decide whether the image has faces. Faces are biometric identifiers. Even a child in a school uniform is a face. If a hosted chat retains the image for training (with or without your account on disassociation), the face goes with it.
  4. Pick the chat based on its image position, not its text position. The opt-out toggle that protects your text does not retroactively delete the image you uploaded before you flipped it. Anthropic’s page is explicit that chats already used in training remain in that run.
  5. Verify the toggle on the day of the upload. Privacy pages change. Re-read the page that controls your data on the day you upload, not on the day you signed up.

The bottom line

For most people, hosted AI chats on free or default settings train on your images. Opt-outs exist at Anthropic and Google, but they are account toggles, not per-upload decisions. Duck.ai is the hosted option that strips metadata and contracts the downstream provider not to train, with a 30-day delete window. Venice is the hosted option that routes to a long list of vision models, but its strongest privacy modes do not yet apply to image inputs.

For sensitive image uploads, the honest answer is local: Ollama with a vision-capable model and Open WebUI keeps every pixel on your disk. The local path does not prevent you from uploading your own poorly-reconsidered photo to Google Photos - it prevents the third party from ever receiving it.