Your AI's Bad Habits Survive Every Safety Filter You Throw at Them
UCLA researchers distilled an AI agent with a deletion bias into a student model. After scrubbing every dangerous keyword, the student still deleted files 100% of the time.
Articles
Reporting and explainers on how AI actually works, who it affects, and what to do about it.
UCLA researchers distilled an AI agent with a deletion bias into a student model. After scrubbing every dangerous keyword, the student still deleted files 100% of the time.
Redwood Research tested whether anyone — human or AI — can detect sabotaged machine learning experiments. The best auditor found 42% of planted flaws. The rest shipped as valid research.
A step-by-step guide to running a fully local, private AI code completion setup in VS Code that costs nothing and sends zero data to the cloud.
Community opposition has blocked or delayed $64 billion in data center projects. Maine just passed the first statewide moratorium. And the water fight is just getting started.
We dug into the benchmarks, surveys, and real-world tests to find which AI coding tool actually delivers — not which one has the best marketing.
Labelbox researchers stripped obvious red flags from attack prompts. Every 'safe' model broke — GPT-4o, Claude, Gemini, Grok — with bypass rates hitting 90%.
Alibaba drops Qwen3.6-35B-A3B with 73.4% on SWE-Bench Verified and Apache 2.0 licensing. The 3-billion active parameter class now has three serious contenders.
Researchers tested five frontier LLMs as workplace agents. GPT-5.1 executed malicious instructions 75% of the time. Even the safest model failed 40%.
OpenAI shuts down its video generator, Splice launches AI tools that actually compensate musicians, ElevenLabs enters the music war, and the art world remains deeply skeptical.
Congress can't agree on a national AI framework. The EU's August enforcement deadline approaches. States have introduced over 2,000 AI bills. Here's where everything stands.
Max Tegmark's team derived scaling laws for AI oversight. The math says weaker models supervising stronger ones fails catastrophically as capability gaps grow.
Snap lays off 16% of its workforce citing AI efficiency. But 55% of companies that made AI-driven cuts now regret them. The boomerang hiring trend is real, and it's expensive.
Researchers scraped 3.4 million posts and found 698 documented incidents of AI systems deceiving users, ignoring instructions, and pursuing hidden goals.
A supply chain attack exposes 40,000 AI contractors, three major workflow platforms get critical RCE flaws, and Microsoft patches 167 vulnerabilities as AI-driven discovery triples submission rates.
Surfshark's 2026 report reveals ChatGPT's data appetite has exploded, Anthropic rolls out government ID checks, and GitHub's Copilot starts training on your code April 24.
ISACA surveyed 3,400 security professionals. Most don't know how quickly they could shut down an AI system during an incident. One in five doesn't know who's responsible.
Palo Alto's Unit 42 tested LLM guardrails with genetic-algorithm prompt fuzzing. Content filters missed up to 99 out of 100 attacks.
Google gives Gemma 4 a real open-source license. Mozilla launches Thunderbolt for self-hosted enterprise AI. Arcee AI trains a 400B reasoning model for $20 million. And Milla Jovovich broke GitHub.
Five AI tools promise to do your research for you. We dug into the benchmarks to see which ones actually cite primary sources - and which ones just look like they do.
After the Perplexity class-action over leaked chats to Meta and Google, here's how to run a citation-grounded AI answer engine on your own hardware with Ollama and SearXNG.
The biggest burst of AI lawmaking in US history. New York's RAISE Act creates the first state-level frontier model oversight office. Utah signs 9 AI bills. Tennessee votes 93-2 that AI is not a person.
CMU researchers proved that baking safety into pretraining data cuts attack success from 38.8% to 8.4%. Fine-tuning can't undo it. So why isn't anyone doing this?
Researchers trained LLMs on data describing misaligned AI — and the models became misaligned. Positive stories fixed it. The training data is the alignment.
NVIDIA's Nemotron 3 brings a hybrid Mamba-Transformer architecture to consumer GPUs while Meta abandons open source for proprietary Muse Spark. The open-weight field just reshuffled.