Deleting a conversation in ChatGPT, Claude, or Gemini removes it from your view, but the message has already left your machine, travelled over a network, and been logged on a server you do not control. The privacy guides on this site cover what happens to data after it lands there (how to delete your data, how to read an AI privacy policy, and the threat model). This page covers the step before any of that: how to redact what you paste so the sensitive parts never leave your machine.
TL;DR
- PII is anything that identifies a person directly or in combination. National ID, account number, full name plus DOB, and biometric records are clearly identifying; gender plus ZIP plus full DOB becomes identifying for the vast majority of a population.
- Three of the biggest AI assistants state in their own policies that inputs may be reviewed by humans or used to train models unless you opt out. Treat the prompt you paste as data the company can read, and redact accordingly.
- Redaction is faster than redaction-by-replacement when you let a local tool do it. Open-source options like Presidio and spaCy run entirely on your machine and can flag names, organizations, locations, IDs, and financial numbers before you send anything.
- Image uploads carry EXIF, GPS, and camera metadata that can include the device, the timestamp, and the location the photo was taken. ExifTool strips those tags locally before upload.
- The honest baseline: automated detectors miss things. Use them to narrow the surface, then read the prompt yourself for anything a tool would not catch - names of small children, internal project codenames, idiosyncratic phrasing that a colleague would recognize.
What counts as PII, by the standard definitions
NIST Special Publication 800-122 defines PII as either “information used to distinguish or trace identity” (name, Social Security number, date and place of birth, mother’s maiden name, biometric records) or “information that is linked or linkable to an individual” (medical, educational, financial, employment information). A user’s IP address alone is not PII, but is “linked PII” once combined with a timestamp or a session.
GDPR Article 4(1) uses a similar but broader test: any information relating to an identified or identifiable natural person, including a name, identification number, location data, online identifier, or one or more factors specific to physical, physiological, genetic, mental, economic, cultural, or social identity. Identification may be direct or indirect.
The combination effect matters. A landmark 1990 study cited in the NIST guidance found that 87 percent of the US population could be uniquely identified using only gender, ZIP code, and full date of birth - three fields most people would not call sensitive. The same math applies to your prompts: even when no single field looks identifying, the conjunction can be. The practical question is not “is each field PII” but “does the prompt point to one person.”
What the AI vendors themselves say happens to what you paste
Three policies, fetched fresh:
- Anthropic. The Privacy Policy defines Inputs as “the content you submit to the Services” and notes that “We may use your Inputs and Outputs to train and improve Anthropic AI models, unless you opt out through your account settings.” Inputs and Outputs are also used “when your conversations are flagged for safety review” or “when you’ve explicitly reported the materials to us,” regardless of the opt-out. Anthropic collects account identifiers (name, email, phone) and technical information (device type, operating system, IP address).
- Google. The Gemini Apps Privacy Hub lists what users say (prompts, spoken input, tasks), what they share (files, videos, screens, photos), Gemini Live transcripts, and generated content as collected data. The page is explicit that “a subset of chats are reviewed by trained reviewers” and that “reviewed data is disconnected from your account before being sent to providers” but is retained for up to three years. Google’s standing guidance: “Please don’t enter confidential information that you wouldn’t want a reviewer to see or Google to use to improve our services.”
- OpenAI. OpenAI documents an opt-out for training use; even with that opt-out, conversations are retained for safety and legal reasons, and human review is part of the abuse pipeline.
The upshot is not “AI assistants leak your data.” Each company documents retention windows and opt-outs. The honest read is: the prompt you paste is text the company can read, and a human reviewer may read it too. Redaction is how you keep the parts you would not want read off the wire entirely.
What to redact before pasting
A pre-flight checklist that fits on one screen:
- Direct identifiers. Full names, home addresses, phone numbers, email addresses, government IDs (passport, driver’s license, Social Security or national insurance), account numbers (bank, credit card), employee or student IDs.
- Indirect identifiers. Date and place of birth, employer name plus role, school plus graduation year, hospital or clinic name plus a date, vehicle plate, IP address combined with a session.
- Contextual identifiers. Unusual phrasing a colleague would recognize, internal project codenames, and references to a specific office or team. These are the categories automated tools will miss.
- Secrets. API keys, OAuth tokens, database connection strings, private SSH keys, signed URLs, internal hostnames, .env contents. These do not look like PII and the redactors below will not flag them. Look for them yourself.
- Metadata on uploads. EXIF (camera, lens, timestamp), GPS (latitude, longitude, altitude), IPTC (caption, byline, copyright), and XMP (edit history, software used). ExifTool reads and writes EXIF, GPS, IPTC, XMP, and MakerNotes; deletion is the “write nothing” operation and runs entirely offline.
A quick read-through before sending is the single highest-yield step. The list above exists to give the read-through structure, not to replace it.
Tools that detect PII locally, before any cloud call
Two open-source tools do the detection entirely on your machine:
- Presidio (Microsoft / Data Privacy Stack) is a fast identification and anonymization pipeline for private entities in text and images. It detects credit card numbers, names, locations, Social Security numbers, US phone numbers, bitcoin wallets, and financial data using Named Entity Recognition, regular expressions, rule-based logic, and checksums. It can run fully on-prem via Python, PySpark, Docker, or Kubernetes. The typical pipeline is: run the Analyzer to identify PII entities, then run the Anonymizer to replace them with operators such as pseudonymization, encryption, redact, or replace. Presidio’s documentation is explicit that automated detection is not a guarantee: “there is no guarantee it will find all sensitive information - additional systems and protections should be employed.”
- spaCy is a Python NLP library with a fast statistical entity recognition system. The default trained pipelines assign labels to contiguous spans of tokens; entity types include
PERSON(e.g., person names),GPE(countries, cities, states),ORG(companies, agencies, institutions),MONEY,DATE, andLANGUAGE. The vectors and pipelines run on CPU by default; GPU is optional. Custom PII patterns (SSN regex, credit card numbers) can be added via rule-based matching.
For image metadata, ExifTool runs as a local Perl script with no network connection required and supports 150-plus file types including images, video, audio, and documents. The standard recipe is exiftool -all= file.jpg to strip all metadata tags before upload.
The point of these tools is not to replace the human read-through. The point is to flag the obvious cases in seconds so the human read-through can focus on the non-obvious ones.
What This Means
Redaction is a habit, not a single action. The minimum viable habit: paste into a local editor first, run a local PII detector over it, read for anything the tool would miss, then paste into the assistant. The marginal cost is a minute per prompt; the marginal benefit is that the most damaging payloads never reach a server in the first place. If a single prompt carries too much identifying context to redact, the answer is usually “use a local model for this one” - not “paste and hope.”
The Bottom Line
Three definitions you can lean on (NIST SP 800-122 for the US, GDPR Article 4(1) for the EU) treat anything that points to a person as personal data, and treat combinations of weak fields as identifying. Three vendor policies (Anthropic, Google, OpenAI) all confirm that inputs may be reviewed by humans or used to train models unless you opt out. Two local open-source tools (Presidio and spaCy) detect the obvious PII without sending anything to a model. One local CLI (ExifTool) strips image metadata before upload. None of this replaces judgment; it raises the floor so judgment can do the rest.