Top Stories
Hackers are stealing Claude tokens from subscribers
A UK consultant noticed his Claude Max 20x account burning tokens while he wasn’t working. Anthropic eventually confirmed to Grant De Swardt that the activity came from a “compromised Claude session key” used to mint unauthorized Claude Code OAuth tokens, then emailed him a warning letter blaming “common infostealer malware.” The malware, the kind that grabs saved passwords and session data from infected machines, had been used to log in to his account and consume usage allowances. Multiple other subscribers reported identical zero-to-100% usage spikes. Anthropic refunded part of the charge and invalidated his sessions and tokens, but TechCrunch reports that the company’s support can only see totals, not itemized usage, so the theft can go unnoticed for months. This is the same harvest-and-resell pattern IBM X-Force tracked earlier this year, when 300,000 ChatGPT credentials surfaced on the dark web after infostealer infections.
This is the cleanest concrete example yet of the “AI stole my data” failure mode our readers worry about. There is no breach of Anthropic’s servers in the conventional sense: the attacker uses your already-valid session token, and the user-facing dashboard gives totals but not the per-request log. The natural test is whether Anthropic eventually ships an itemized usage feed and a way to revoke Claude Code OAuth tokens on demand. De Swardt’s response: he cancelled his Claude subscription and moved to Cursor, which can route through multiple model providers. The De Swardt incident is also the highest-traffic privacy story of the week, and the failure mode here is structural, not a one-off.
OpenAI says it solved the Navier-Stokes Millennium Problem; academics say it scooped them
OpenAI’s Sébastien Bubeck and Mark Chen announced on September 8 that roughly 10,000 concurrent OpenAI agents had resolved the Navier-Stokes existence-and-smoothness problem “at a cost of millions.” Simon Willison’s deep dive puts the total API cost at an estimated $15 million across all attempted problems (4.9 million messages and about 300 billion output tokens at GPT-6 Astra public prices); the Navier-Stokes resolution alone used 2.7 million messages and roughly 130 billion output tokens. Agents ran for about 88 hours before a 17-hour Lean verification pass. NYU’s Tristan Buckmaster, who had been working on a related simplified-Euler version with Anthropic’s Levent Alpöge for almost a year and posted a proof of his own, says OpenAI staff offered to publish first or jointly without Alpöge due to his Anthropic affiliation. Buckmaster asked whether the agents had read or been trained on transcripts of his and Alpöge’s Codex sessions; he was told the model did not look up user data and received no answer on training. Mark Chen “again denied” transcript access.
Terence Tao, quoted the next day by Simon Willison, warned that the rumor of a problem being worked on can now trigger “a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential,” and that “prematurely solving the problem by purely AI-powered methods…can contaminate this process to the point where it actually becomes a net negative.” MIT Tech Review’s quoted researchers flagged the same incentive problem: a few well-funded labs can swallow problems whole, and “very few mathematicians will have resources of that scale.” This is the first public proof-of-scooping dispute between a frontier lab and named academic mathematicians; the transcript-access question is going to keep coming up.
Meta launches Muse, a personal AI agent that wants your email, calendar, payments, and health data
Meta unveiled Muse on September 8, its biggest consumer AI bet yet and the first major launch two weeks after the company’s $18 billion multistate settlement over youth harm on its platforms. Muse is a personal AI agent that runs in a “dedicated Muse Secure VM” with a separate Sentinel agent, can opt in to read a user’s email, calendars, payments, health and fitness apps, smart home, dining, shopping, music, and events, and uses Meta’s Muse Spark model. Pricing tiers at launch are a free tier (with a usage meter), $20/month Power, and $100/month Maximum, and it’s available on the web, iOS, Android, and WhatsApp. Meta claims Muse “doesn’t share people’s conversations or data with Meta’s ads systems” and uses Link by Stripe for checkout, but the technical details are at research.meta.ai and “require independent security verification,” in the article’s framing.
The load-bearing context is Meta’s privacy track record: the 2011 FTC consent decree, the 2019 $5 billion FTC penalty (plus the same-year discovery of “users’ passwords in readable formats”), the 2023 FTC charge that Meta violated the 2019 order, Cambridge Analytica, and the August 2026 $18B settlement. None of that history appears in the announcement, and Meta’s privacy claims are not externally audited. For intelligibberish readers, the practical questions are whether Muse actually isolates ad tracking from agent data at the OS/VM level and whether Meta publishes the audit. The story is set up for either a privacy-trust walkthrough or a local-AI counterpoint.
Mistral raises €3B Series D at a €21B+ valuation, framing itself as Europe’s “third way” in AI
Mistral closed €3 billion on September 8 at a post-money valuation of more than €21 billion. Samsung Electronics led; EQT’s Scaleup Europe Fund and existing investor PSG Equity co-led; a16z, Nvidia, and Salesforce Ventures participated, with Advent, BlackRock, and the Grand Duchy of Luxembourg as new backers. Mistral called it “the largest equity fundraising round ever completed by a European technology company.” Capital is earmarked for 1 GW of European compute capacity by 2030 and “international growth”; French president Macron publicly framed the round as France and South Korea building “a third way in AI” between the US and China.
For local-AI readers the angle is real even if Mistral stays mostly closed-weight: the company already hosts third-party open-weight models on its platform and in August rolled out region selection for AI query processing, both of which map to the “self-hosted AI” and “sovereign AI” demand seeds we track. ASML (a strategic partner and investor) and the expanded Microsoft partnership from July are the supporting hardware and cloud plumbing.
Cognition hits a $48B valuation as investors double down on non-winner-take-all AI coding
Cognition closed a $2 billion Series E at a $48 billion valuation, four months after a $26 billion round in May. Andreessen Horowitz, Accel, Founders Fund, General Catalyst, and Avenir led. Run-rate revenue moved from $492 million in May to $900 million now and is projected at $4-5 billion by year-end, per the report. Cognition is also training its own model based on open-source alternatives to reduce reliance on OpenAI and Anthropic, and leases an Nvidia server cluster that costs hundreds of millions a year, with total cash burn potentially reaching $800 million this year (per The Information). Customers include Mercedes-Benz, NASA, Goldman Sachs, and Citi.
The investment thesis in TechCrunch’s framing: coding is not a winner-take-all market, and Cursor’s earlier $50B-in-talks valuation (April) and eventual $60B sale to SpaceX were not the end of the category. a16z led both rounds, which is itself a signal that top-tier investors want exposure to multiple coding agents rather than picking one.
EFF: federal judge rules DOD’s “supply chain risk” tag on Anthropic was unlawful First Amendment retaliation
EFF reported on September 1 that a federal judge ruled the Department of Defense’s “supply chain risk” designation of Anthropic was unlawful retaliation in violation of the First Amendment. The designation followed Anthropic’s refusal to let Claude power mass surveillance of Americans or autonomous weapons. Matthew Guariglia’s post notes that EFF and allies filed amicus briefs arguing the Pentagon’s action violated Anthropic’s First Amendment rights, and frames the holding as “the government cannot punish a company for having preferences regarding unconstitutional uses of its technology.” EFF flags the limits of the win: “We shouldn’t have to rely on private companies to protect us from the surveillance state. It’s past time for Congress to act.”
This is the freshest court ruling tying a vendor’s terms-of-use decisions to First Amendment protection, and the EFF framing leaves open whether a company’s choices about how its technology may be used are protected speech in their own right. For privacy: the precedent strengthens the position of any frontier lab that refuses a specific government use case, but it does not bind Congress. We covered the underlying fight in our Anthropic v. Pentagon writeup when Judge Rita Lin first blocked the designation.
OpenAI ships ChatGPT Images 2.5 with two new API models, Sunburst and Flare
OpenAI released ChatGPT Images 2.5 on September 8, adding two new API models under the IDs gpt-image-2.5-sunburst and gpt-image-2.5-flare. OpenAI’s announcement positions Sunburst for editing precision and Flare for “fast, high-quality everyday image generation,” with better multi-turn instruction following and improved preservation of subjects across reference photos. OpenAI notes that its image models have now produced more than 3 billion images across ChatGPT Images and the GPT-Image API. Simon Willison’s CLI demo shows Sunburst adding a raccoon scientist to an existing chart while keeping the chart’s structure intact, which is the multi-turn editing case in plain view.
404 Media: a secretive DHS “predictive policing” unit is scoring Americans’ financial habits and ordering traffic stops
404 Media reported on September 8 that U.S. Border Patrol predictive policing units analyze Americans’ financial activity and other data, then feed that intelligence to local police who pull over individuals “not suspected of any specific crime.” Joseph Cox’s piece describes the units as secretive, with no specific victims or unit names disclosed in the accessible portion. The story fits the pattern we have covered before: AI-on-civilians risk scoring is migrating from local police departments into federal agencies that share data with them.
Quick Hits
- Flock taught cops how to surveil “No Kings” protesters. 404 Media reports that a Flock webinar (“Prepared For Anything: How Cities Prepare For Planned And Unplanned Events”), hosted by Flock’s director of market management Caity Peak, walked agencies through combining ALPRs, drones, gunshot detectors, 911 data, and live video feeds to monitor protests. Peak’s stated goal: monitor protests “without having cops physically on the ground ‘making it worse, getting attention.’” Flock now claims 4,800+ cities run FlockOS. The pattern matches our earlier look at the LAPD ALPR audit, which found 161 false stolen-car stops in 60 days from Flock’s cameras before the contract lapsed.
- Chrome now ships updates every two weeks. TechCrunch reports Chrome 153 launched on desktop, iOS, and Android on a two-week cadence, down from the four-week cycle adopted in 2021 (itself down from six weeks). Google attributed the acceleration to AI-enabled threats and to higher patch volume driven by automated AI tools and community bug reports. Mozilla, Microsoft, and Brave have also moved to two weeks.
- Google Cloud signs an Accenture deal to catch up in enterprise AI deployment. TechCrunch reports the new Accenture Gemini Enterprise Business Group, which will train up to 1,000 Accenture forward-deployed engineers to build on Google’s Gemini Enterprise platform. Google’s Q2 revenue was $24.8 billion; per Ramp’s August 2026 data, Google’s share of enterprise AI spend among U.S. customers was about 6%, behind Anthropic (43.5%) and OpenAI (39.7%).
- Hugging Face: “Safety for Whom? Refusing the right subset of a topic, not the whole topic.” Multiverse Computing CAI argues that safety alignment should treat harm as a property of a subset within a topic, not of the topic itself. On Qwen3-8B, the approach raised political refusal from 9.47% to 84.75% and cut mean unsafe-response rate across HarmBench, StrongREJECT, and WildJailbreak from 26.26% to 0.14%, at the cost of XSTest over-refusal rising from 2.00% to 74.00% in the strongest config. Posted September 8, paper ID 2609.04482.
- Hugging Face: “NeoMME: an efficient multimodal-native and multilingual encoder.” H Company introduces a 260M/800M-parameter encoder that processes text tokens and 32x32 image patches through a single bidirectional Transformer, without a separate vision tower. The 260M model encodes roughly 51 pages/sec at 2048x2048 on an NVIDIA L40S (about 2x ColModernVBERT) and reduces late-interaction index storage from about 1.5 MB to 6 kB per page. Posted September 3, Apache 2.0.
Worth Watching
- Whether Anthropic ships an itemized Claude usage feed. The current “totals only” view is the structural failure that lets the session-key theft stay invisible for months; an itemized feed plus OAuth-token revocation would close the loop.
- Whether the Navier-Stokes dispute produces a transcript-access disclosure rule. Buckmaster asked for, and did not get, an answer on training-data usage; the next test is whether OpenAI, Anthropic, or any lab publishes a policy for that class of question.
- Whether Meta publishes the Muse privacy audit. The “dedicated Muse Secure VM” and Sentinel agent architecture are testable claims; the natural test is an external audit or red-team report on whether agent data is actually isolated from the ads stack.
- Whether Mistral’s 1 GW target holds to 2030. Samsung’s lead is the load-bearing commitment; the next signal is whether Mistral announces a chip supply or datacenter partnership of comparable scale.
- Whether the DHS predictive policing unit story produces a congressional hearing. Border Patrol financial-data scoring and “pull over without a specific crime” traffic stops are a textbook AI-and-civil-liberties story; the test is whether any committee demands the unit name and the data sources.
- Whether Google’s Accenture FDEs convert into measurable Gemini Enterprise share. Ramp’s August 2026 numbers have Google at 6% vs Anthropic’s 43.5%; the test is the December cut of the same dataset.