How to Read an AI Privacy Policy Before You Sign Up

The five sections that actually matter in any AI privacy policy, what each major assistant's policy really says, and the red flags that mean walk away.

A privacy policy is a legal document, not a commitment. The wall of text you scroll past to click “I agree” exists so the company can argue in court that they told you. Reading it for trust, rather than for legal exposure, takes about five minutes if you know the five sections that matter and the phrases that mean walk away. This guide walks through both, then compares what five real AI vendors actually say so you can see the difference between a policy that respects you and one that respects its own lawyers.

TL;DR

  • A privacy policy is six things. Identity and contact details of the controller, what they collect, why they collect it, who they share it with, how long they keep it, and your rights. Anything else is padding.
  • Training is the question, not retention. The retention number is the same across the industry (30 days for backups) because that is the standard backup cycle. What differs is whether your future conversations train the model.
  • The opt-out toggle is not the end of the story. Anthropic, Google, and OpenAI all spell out carve-outs where they will still use flagged or feedback-flagged content for model improvement even if you opted out.
  • Duck.ai and Proton Lumo take a stronger position: contractually binding providers to no-training and using zero-access encryption for stored chats. Duck.ai strips your IP address before the request ever reaches the model provider.
  • Under GDPR Article 13, the controller must tell you the retention period, the recipients, the legal basis, and the right to withdraw consent, at the time of collection. If a policy is missing any of these for your region, that is itself a red flag.

What a privacy policy is actually for

The point of a privacy policy is not to inform you. The point is to give the company a defense if a regulator, a journalist, or a plaintiff asks “did you tell them?” Decades of US FTC consent decrees and GDPR enforcement (in force since May 2018) have pushed the legal industry toward a single template: every required disclosure, buried in the order regulators expect, in language that no reader will wade through. Reading it for the actual answer to “what happens to my data” is a different exercise.

You are not reading for trust. You are reading for the five answers you need before you type a prompt.

The five sections that actually matter

A well-formed policy, whether it is 12 pages or 60, has the same shape. Cross-reference to the GDPR Article 13 list and you will see why: regulators tell controllers they must publish these items, so vendors publish them.

1. Identity and contact details

GDPR Article 13(1) requires the controller to publish its identity, contact details, and (where applicable) the data protection officer. Anthropic lists a [email protected] address and a separate [email protected] for the DPO. Mistral publishes a retention help article on help.mistral.ai. If a privacy policy has no named controller, no physical address, and no contact channel that does not route through a generic form, walk away: there is no one to write to when something goes wrong.

2. What they collect

Every AI privacy policy collects three buckets: what you typed (Inputs), what the model replied with (Outputs), and the technical metadata around it (IP address, device info, timestamps). The exact categories are spelled out in each company’s “Data We Collect” section. Anthropic enumerates Identity and Contact Data, Payment Information, Inputs and Outputs, Feedback, Study Participation Data, Communication Information, Verification Data (which may include biometric data), and Technical Information. Google’s Gemini Apps Privacy Hub lists prompts, files you share, transcripts of Gemini Live audio, feedback, Generated content, model thinking steps, Connected Apps data, device info, general-area location, and “Remote browser data like cookies that contain your website authentication info.”

The eyebrows go up on the categories you did not expect. Anthropic’s “Verification Data (may include biometric data)” line means the system may collect face or voice prints during identity verification, depending on the path you took to sign up. Google’s “Remote browser data like cookies that contain your website authentication info” line is what makes the Gemini browser extension a different privacy object from the chat page.

3. How long they keep it

The retention number that gets quoted in the press is rarely the retention number that matters. Anthropic says deleted chats are “removed immediately from your conversation history and automatically deleted from our back-end within 30 days.” Google’s default for Gemini Apps is auto-delete after 18 months, adjustable to 3 or 36 months or off, and 72 hours when Keep Activity is off. Mistral’s help article gives specific numbers: civil identity data kept for 5 years after the end of the user contract, other account creation data for 1 year after account deletion, technical data for a rolling year, and invoices for 10 years from the close of the relevant financial year.

Two layers sit underneath that headline number and matter more:

  • Backups. The 30-day figure that appears in every major policy is the standard backup cycle. The chat you deleted on Monday may live in a backup snapshot until the cycle rolls forward.
  • Human-review copies. Google’s Gemini Apps Privacy Hub is explicit: human-reviewed chats are “retained for up to three years” and “disconnected from your account before being sent to service providers.” Anthropic’s policy says flagged or feedback-flagged content is retained on a separate track for trust and safety model training.

4. Training opt-out, with the carve-outs

This is the section that changes between vendors and the one that journalists most often get wrong. The toggle is the easy part; the carve-outs are where the actual training happens.

  • Anthropic. Per the Privacy Policy: “We may use your Inputs and Outputs to train and improve Anthropic AI models, unless you opt out through your account settings.” Even if you opt out, the policy says Anthropic will use your inputs and outputs “when your conversations are flagged for safety review to improve our ability to detect harmful content, enforce our policies, or advance AI safety research” or “you’ve explicitly reported the materials to us (for example via our feedback mechanisms).”
  • Google Gemini. The Privacy Hub is direct: when Keep Activity is on, “Google uses your activity to provide, develop, and improve its services (including training generative AI models).” Turn Keep Activity off and “your future chats won’t be used to train our AI models, unless you choose to send Google feedback.” Even then: “Google still uses your chats to respond to you and help protect Google, our users, and the public, including with help from human reviewers.”
  • OpenAI ChatGPT. For the API, OpenAI’s “Your data” documentation states: “As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us).” For ChatGPT consumer accounts, a separate opt-out toggle lives under Settings - Data Controls - “Improve model for everyone.” The general rule the documentation applies is that toggling off affects future data only; past chats already in a training set are not retroactively removed.
  • DuckDuckGo AI Chat (Duck.ai). The Duck.ai Privacy & Terms take a different position: “we have agreements in place with all model providers that further limit how they can use data from these anonymous requests, including not using Prompts and Outputs to develop or improve their models.” Providers named on the privacy page include Anthropic, Azure OpenAI, OpenAI, and together.ai. The contract is the privacy promise.
  • Proton Lumo. The Lumo security model page frames the position architecturally: “the request is not logged or retained by the LLM server after completing that request” and “the LLM server forgets the request as soon as the response is generated.” Combined with zero-access encryption for stored chat history, the claim is that no party in the chain keeps a usable copy.

The pattern: a vendor that owns the model gives you a toggle and a list of carve-outs. A vendor that proxies to someone else’s model can promise by contract. A vendor that runs the model in its own zero-data-retention infrastructure can promise by architecture.

5. Sharing and recipients

GDPR Article 13(1)(e) requires the policy to list recipients or categories of recipients. Three patterns show up:

  • Affiliates and subprocessors. Anthropic lists “Affiliates and corporate partners,” “Service providers and business partners,” and “Your Organization and Administrator,” plus a separate subprocessors list at anthropic.com/subprocessors.
  • Connected apps and “agents.” Google’s policy explains that Gemini will share “with other services and third parties necessary info” to complete tasks you ask it to perform, which is the section that matters most if you let the assistant browse, send mail, or write to your calendar.
  • Government requests. Every major policy reserves the right to disclose to “government authorities, law enforcement, or other third parties” for legal or safety reasons. The published transparency reports are the proof of how often that right is exercised.

Worked example: how five policies answer the same question

Reading the same five sections across five vendors is the fastest way to see the difference.

QuestionAnthropicGoogle GeminiOpenAI ChatGPTDuck.aiProton Lumo
Train on your future chats?Yes, opt-out availableYes if Keep Activity onYes, opt-out availableNo (contractual)No (architecture)
Opt-out applies to past data?No (already-trained stays)No (retained)No (already-trained stays)N/AN/A
Human-review retentionTrust and safety models, de-identifiedUp to 3 years, account-disconnectedUp to 30 days backups plus legal holdsAt most 30 daysNot retained
IP exposed to upstream?N/A (runs own model)N/A (runs own model)N/A (runs own model)No (proxy strips)N/A (runs own model)
Stored chats encrypted from vendor?No (vendor can read)No (vendor can read)No (vendor can read)Local-first; sync deleted after 18 months idleYes (zero-access encryption)

The Duck.ai and Lumo columns are not the same product (Duck.ai routes your prompt to a third-party model under contract; Lumo runs its own model on its own hardware), but they answer the same question - “does the vendor train on me?” - in the same direction.

What the law actually requires

GDPR Article 13 is the cleanest list of what a controller must tell you. At the time of collection, the policy must include the controller’s identity, the DPO’s contact details where applicable, the purposes and legal basis of the processing, the legitimate interests where Article 6(1)(f) applies, the recipients, the international transfers, the retention period or the criteria for setting it, your rights (access, rectification, erasure, restriction, objection, portability), the right to withdraw consent, the right to lodge a complaint with a supervisory authority, and information about automated decision-making under Article 22. If a policy published for EU users is missing any of these, that is itself a finding you can raise with the supervisory authority in your country.

Outside the EU, the picture is uneven. The US has no general federal privacy law and no general right to erasure. California (CCPA/CPRA) gives residents the right to delete personal information, with business-disclosed exceptions for security, fraud prevention, and legal compliance. Other US states have similar but distinct rules.

Red flags that mean walk away

Five phrases and one absence that mean the policy does not protect you:

  1. “May include” in the data-collected list with no bound. “Verification Data (may include biometric data)” is an example of bound language. “We may collect any information we deem necessary” is not.
  2. “Service providers and business partners” with no list. Subprocessors are a list, not a vibe. If anthropic.com/subprocessors or the equivalent is empty, ask why.
  3. “We may share with affiliates and other entities” with no categories. GDPR requires categories, not vibes.
  4. “As needed for legitimate interests” with no specification. Article 6(1)(f) lets controllers process on legitimate-interest grounds, but the legitimate interest must be named.
  5. “We may modify this policy at any time” with no notice mechanism. A policy you cannot track is a policy you cannot hold anyone to.
  6. No retention number, only “as long as necessary.” Article 13(2)(a) requires the retention period or, if impossible, the criteria used to determine it. “As long as necessary” alone fails this test.

The Bottom Line

A privacy policy is not a promise; it is a leak map. Read the five sections - who, what, how long, training, sharing - and you have 90% of what you need to decide whether to type. The retention number is the same across the industry because backups are the same across the industry. The training opt-out is the same in name across vendors and very different in substance: Anthropic and Google and OpenAI give you a toggle plus carve-outs for safety review and feedback; Duck.ai and Lumo make no-training a contract or an architecture, not a setting. The law gives EU users a specific checklist (GDPR Article 13); outside the EU, the same checklist is the standard to demand. If the policy cannot name its controller, list its subprocessors, give you a retention number, or describe the training carve-outs, the product is not better than the policy that fails to describe it.