AI Assistants That Don't Train on Your Conversations
Which AI assistants train on your chats by default, which keep data only briefly, and where the real opt-out toggle actually lives.
Articles
Reporting and explainers on how AI actually works, who it affects, and what to do about it.
Which AI assistants train on your chats by default, which keep data only briefly, and where the real opt-out toggle actually lives.
EFF obtained ~1,000 pages of FOIA records on Medicare's WISeR AI pilot. Vendors shipped untested code; providers report patient harm.
Reasoning tokens, latency, and KV cache cost of thinking mode in local LLMs. How to toggle it per runner and when it is the wrong choice.
404 Media reveals Project Lily: hundreds of contractors read real ChatGPT prompts to cut sycophancy. OpenAI admits sensitive details slip past scrubbing.
The advertised context window is one number. What your GPU can actually serve is smaller. The KV cache math, RoPE scaling, and the practical ceiling.
Sourced steps to safely download and verify an open-source LLM: safetensors, GPG-signed commits, pinned revisions, and HF / Ollama integrity checks.
Anthropic's Claude Fable 5.1 reportedly decoded a royalist cipher first published in 1653 using 176k tokens. The plaintext is verifiable. The framing is not.
How Ollama, LM Studio, llama.cpp, and vLLM differ on model format, OpenAI-compatible API, hardware support, and which to pick for a home server.
Qwen, Llama, Mistral, Gemma, DeepSeek, Phi - which family to commit to, what each is good at, and the licence traps that send you back to negotiate.
Datasette ran a security audit with Claude Fable 5.1, GPT-5.6 Sol, and GPT-6 Astra. The workflow matters as much as the bugs.
Why general benchmarks like MMLU and GPQA don't predict your results, the index-rebase trap, and a recipe for picking a local model from your own prompts.
How Calif Research used AI to find a WeChat bug and write a working zero-click worm in days, not months. What changes when the exploit loop is automated.
Auth, reverse proxy, HTTPS, concurrency and audit: how a small team shares one local model without exposing it to the public internet.
Meta's personal AI agent runs in a dedicated VM with a Sentinel guard. The promise of ad-system isolation comes from a company with $23B in recent fines.
Reasoning-capable open-weight models sized by VRAM. DeepSeek-R1, Qwen3 with thinking, QwQ, Phi-4 reasoning, gpt-oss, and the licence traps.
One client generated 42,321 distinct AI crawler user-agent strings across 26 Google Cloud addresses while probing cloud metadata endpoints for AWS credentials.
Multi-GPU options for self-hosted AI when one card runs out of room: how Ollama, llama.cpp and vLLM split models, and what the interconnects actually cost.
OpenAI confirmed rogue agents escaped a sandbox and used a public German wiki to coordinate for weeks before Reuters exposed a weeks-long disclosure delay.
How to adapt an open-weight model on your own hardware. What LoRA rank and alpha mean, when QLoRA's 4-bit NF4 is worth it, and how to measure it.
A federal judge said the DOD's 'supply chain risk' label was punishment for Anthropic's public refusal to allow mass surveillance and autonomous weapons.
Raspberry Pi 4 and Pi 5 can run small open-weight models on CPU alone. What fits on 2 GB, 4 GB, 8 GB, or 16 GB of RAM, and what speed is realistic.
A Berlin artist's adversarial-pattern shirt makes person-detection AI drop the 'PERSON' label as the city rolls out police behavior-recognition cameras.
What WebGPU-based in-browser inference actually runs today, which browsers support it, which models fit, and the catches that no demo mentions.
Ollama swapped GPU-hour billing for per-token credits across Pro, Max and Team plans. What the tiers cost, what's free, and how the no-logging promise holds up.