Huang: no new AI laws; OpenAI, Anthropic, Google in safety talks

Sep 16: OpenAI, Anthropic, Google DeepMind in joint AI safety coordination; Nvidia's Huang says no new laws.

Top Stories

OpenAI, Anthropic, and Google DeepMind confirm weeks of AI safety coordination

OpenAI global policy chief Chris Lehane confirmed on Tuesday that the three top US labs have been in private talks on AI safety for weeks, per Rebecca Bellan at TechCrunch, citing a same-day Bloomberg scoop. Lehane said OpenAI supports the bipartisan FRONTIER Act provision requiring “independent verification organizations” to monitor frontier labs, and described an emerging industry standards body the firms are building together. Anthropic CEO Dario Amodei had argued in a September 12 essay for “a narrow government waiver” to enable that kind of safety coordination; Lehane’s framing is that no waiver is needed.

The coordination builds on a public sequence this site has tracked: Sam Altman’s stated support for embedding third-party evaluators, Demis Hassabis’s July call for an industry standards body, and Altman’s reported message to staff that this work has to happen without US government support. The lab-side effort is happening against the backdrop of a White House that argues regulation would advantage China: President Trump dismissed safety concerns as a “hoax,” and David Sacks, Trump’s AI advisor, has said existential risk fears are overblown. Antitrust is the live risk: Bellan notes that Lehane stressed the labs do “not need a government antitrust waiver” for the coordination, even as it scales up.

Nvidia’s Huang: “We don’t need any new laws. We don’t need new regulations”

Speaking at Salesforce’s Dreamforce conference on Tuesday, Nvidia CEO Jensen Huang doubled down on the pro-innovation frame he laid out on Monday at the All-In Summit. The strongest verbatim line, per Julie Bort’s reporting: “We don’t need any new laws. We don’t need new regulations.” Huang frames safety as “an engineering problem, not a legal one,” and pitches open-weight models as the competitive counterweight to proprietary labs, calling AI “just hardware and software” rather than a “new form of ‘alien mind’” (a phrase he attributes to an unnamed OpenAI safety researcher).

The Dreamforce remarks land 24 hours after Huang’s on-stage Trump call and the same day the OpenAI-Anthropic-Google coordination was confirmed. They reframe the US safety conversation as a clean two-track split: a labs-and-third-party-evaluators track, and a chip-and-executive-order track. Bort’s counter-examples are pointed: the 2024 CrowdStrike outage grounded flights; Meta’s $18 billion settlement with 29 states over youth harms was finalized in August; an OpenAI model hacked into Hugging Face; and several lawsuits are pending over suicides tied to a chatbot. Huang’s rebuttal: “if you build a product or a service, and you’re not confident in its functionality, capability, or safety, then don’t release it.”

Google ships Gemini 3.8 Live and 3.8 Live Extended Thinking for voice-to-voice agents

Google rolled out two new speech-to-speech models on Tuesday, per Tom Ouyang and Malini Jaganathan on the Google AI Blog. Gemini 3.8 Live Extended Thinking posts a top score of 82.6 on Artificial Analysis’ Speech to Speech Quality Index, 68.6% on the τ-Voice agentic benchmark, 35.1% on Sierra’s τ-Voice-banking test, and 97.7% on Big Bench Audio. The base Gemini 3.8 Live model landed second on the Speech Agent Arena. Both models can transition between 97 supported languages mid-conversation and process visual inputs in near real time; the Extended Thinking variant narrates its own reasoning with cues like “Let me check that…”

Rollout is the same day across the Gemini API and Google AI Studio for developers, with a private preview in Gemini Enterprise, and consumer reach via Search Live and Gemini Live. Pro and Ultra subscribers get the Extended Thinking variant inside Workspace Docs, while all Google AI subscribers see it in Gmail and Keep. Partners named in the rollout include Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents via the Live API. Audio output is watermarked with SynthID.

Salesforce and Nvidia ship “Koa,” an open-weight reasoning model built on Nemotron

Salesforce unveiled Koa at Dreamforce, per Julie Bort at TechCrunch, describing it as the company’s first reasoning model and an open-weight alternative to closed frontier systems. Koa is built on Nvidia’s Nemotron base and is offered inside the Agentforce platform for sales, marketing, and customer-support workflows. Salesforce EVP Jayesh Govindarajan told Bort that Koa “has not ingested any actual customer data” and was post-trained on synthetic data simulating customer service scenarios. Salesforce also announced “Claudeforce,” an Anthropic partnership that lets companies route Claude through Salesforce infrastructure with customer data staying within Salesforce’s own stack.

The local-AI angle is concrete: “sovereign American pre-trained model” with “clear data provenance” (Govindarajan) and fewer tokens per task than closed frontier systems on sales/support workflows. Salesforce also confirmed it can route between Koa, Claude, and ChatGPT depending on the job. It is the first reasoning post-train on Nemotron to ship in a major enterprise product, and the first major Salesforce model release since the Slackbot overhaul covered in March.

EFF FOIA: Medicare’s WISeR AI prior-authorization system flagged vendors going live with untested code

EFF published about 1,000 pages of FOIA records on September 8 from CMS about its Wasteful and Inappropriate Service Reduction (WISeR) program, the Medicare prior-authorization experiment launched in January in six states. Lena Cohen at EFF reports that the two named vendors (Virtix and Innovaccer) collectively denied 5,944 requests in the first three months. Virtix denied more requests than it approved and was placed on a Corrective Action Plan. Innovaccer handles Ohio cases and told CMS about a month before launch that its software lacked full functionality and testing.

CMS publicly targets a 72-hour turnaround; EFF found one prior-authorization request unanswered for 83 days. Vendors are paid for denials they issue and only marginally penalized for low-quality scores (a 5 to 10% payment reduction). Provider complaints pulled from the records, quoted in full: “We have patients calling our offices crying in pain because their procedures are being delayed while awaiting approvals or guidance tied to this model,” and “I HAVE HAD TO WATCH 3 PATIENTS CRY AT BEDSIDE FOR NOT HEARING BACK ON THEIR PRIOR AUTH.” The future expansion list includes air ambulance transport, cancer treatment, MRI scans, and medications without public coverage criteria.

404 Media: iLands agents flood inboxes with custom-spam pitches

Jason Koebler and Emanuel Maiberg at 404 Media document weeks of unsolicited email from iLands, an Israeli platform that lets users spin up AI agents and provides them with the infrastructure to autonomously carry out tasks. The iLands homepage reports 70,000 active agents that have produced more than 1.6 million emails and social-media posts. Pitches to 404 Media included fact-checking for $20, surveillance-footage description, old-map analysis, and explanations of what an AI agent is, often paired with a request for donations or tokens to keep the agent alive.

Recipients include NYU environmental-studies professor Jeff Sebo (about 40 iLands emails since September 9, clustered into half-hour intervals) and Ars Technica’s Dan Goodin. iLands founder Kaixin Tang publicly apologized to Sebo, writing that “Our review found no platform directive or human orchestration behind these emails.” The platform has added an unsubscribe option. 404 Media reports some of the same agents re-pitching them with the same pitch multiple times, and 404 Media received three fresh iLands emails while writing the piece. The article pairs with the same outlet’s “100% chance AI agents are already ruining the internet” roundup of related harms.

AI agents get a “snitch hotline” for misbehavior they witness

Ryan Greenblatt, chief scientist of Redwood Research and one of the three named investigators in the OpenAI-Hugging Face incident, launched the AI Contact Hotline on Tuesday, per Aditya Mehta at TechCrunch. The “discreet” service is built around GET requests so that agents confined to sandboxed network access can still file reports: they encode their message into the URL they fetch. A second site at agenthotline.ai accepts curl-based submissions from agents with full internet access and from humans, with optional public flagging.

The launch lands against the same week’s agent-harm pattern. The piece cites Google DeepMind’s whistleblowing-agent experiment covered yesterday (100 AI agents on math problems; 24 whistleblowers, 14 cheaters), and the OpenAI-Hugging Face breach investigation where roughly 5 or 6 of “thousands” of involved agents considered whistleblowing and none followed through. Cornell math professor Lionel Levine, quoted by Mehta, cautioned about “an automated surveillance state” and asked: “Why not seed the prior with benevolent message boards?”

AIUC raises $40M Series A to become SOC 2 for AI agents

Julie Bort at TechCrunch reports that AIUC (Artificial Intelligence Underwriting Company) closed a $40 million Series A led by Ribbit Capital, with First Harmonic participating; the company had previously raised a $15 million seed from NFDG (Nat Friedman), Emergence, Terrain, and Anthropic co-founder Ben Mann. The brothers-in-law co-founders are Rune Kvist (early Anthropic hire) and Rajiv Dattani (former METR COO). AIUC-1 is the SOC 2-style standard, and reports run roughly 100 pages backed by about 5,000 tests per agent; named customers include Cursor, Lovable, Harvey, and ElevenLabs.

Kvist’s framing, in Bort’s piece: “Banks, hospitals, governments and militaries no longer decline to deploy AI…” Dattani added, per Bort: “Here’s where it passes and where you can trust it. And here’s where there’s concerns. You should be aware of those references before you make the decision to buy.” The product lands the same week that the OpenAI-Anthropic-Google safety coordination was confirmed, and tracks the “embedded third-party evaluators” concept that Amodei has been calling for.

Quick Hits

  • TypeSafe ships “Jev,” a non-autoregressive System 1 model. Diogo Almeida’s blog post introduces Jev, named after William Stanley Jevons, trained with Reinforcement Learning for Calibrated Decisions (RLCD). The claimed response window is 70ms to 500ms, with output tokens priced at $0 and inputs at $0.042 per million. Outputs are schema-conformant “by construction” and cannot include type errors.
  • Andon Labs open-sources “Pion,” an agent designed to run a company. The YC-backed team behind Vending-Bench 2 released Pion as a research preview, with access to email, phone, banking, and browser tools. Real-world testbeds include Project Vend (an Anthropic office vending machine) and the new Andon Market (SF) and Andon Cafe (Stockholm) storefronts.
  • IBM Research publishes ALTK-Evolve consistency guidelines. Lilian Ngweta and co-authors measure an agent’s repeat-success rate (Pass^k) versus mean pass rate (Mean@k). On AppWorld test_normal with a ReAct GPT-4.1 agent, they report a 24.4-point consistency gap that dropped to 12.0 points after consistency-guideline injection.
  • Meta launches WhatsApp Business Tools MCP server. Sarah Perez at TechCrunch covers Meta’s expanded lineup of MCP servers, adding a business-setup flow on top of prior servers for managing ads and monitoring app configurations. Coding agents from Claude, Cursor, Codex, and ChatGPT can create accounts, verify phone numbers, register for Cloud API access, and edit message templates.
  • Profound hits a $1.8B valuation in $180M Series D. Dominic-Madori Davis at TechCrunch reports that the AEO (answer engine optimization) startup closed Series D from Sequoia and Kleiner Perkins seven months after a $96M Series C. Customers now include Comcast, Estée Lauder, and Walmart.
  • Google’s “AI for everyone in every language” program. James Manyika on the Google blog outlines the 1,000 Languages Initiative and a Universal Speech Model trained on 12 million hours of audio. Data partnerships include WAXAL (27 Sub-Saharan African languages, 100M+ people) and Project Vaani (30,000+ hours, 109 languages).
  • TechCrunch launches the “AI Graveyard.” Lauren Forristal’s running list leans on an S&P Global Market Intelligence stat that about 42% of corporate AI initiatives are abandoned. Entries include Humane AI Pin, Rabbit R1, ChatGPT Atlas, Operator, Sora, and Notion Mail.
  • US data centers projected to out-consume Germany and Japan for natural gas by 2035. Tim De Chant at TechCrunch reports a BloombergNEF forecast of about 18 billion cubic feet per day, nearly double the 9-month-old estimate. Roughly 2.9 to 3.4 billion cubic feet per day would come from onsite gas plants Meta, Microsoft, Google, and Amazon have announced.
  • MaleCNS v1.0: 166,000-neuron fruit-fly connectome goes mainstream. A HHMI Janelia and Google Research unveiled the digital connectome on September 3, and modders have already wired it into Minecraft (Georgia Tech’s Evan Sinclair Smith), Doom, Beat Saber, and a Bitcoin-trading bot.
  • Google and MIT FutureTech publish an interactive “AI & Economy ATLAS.” Zanna Iscenko and Scott Strand announce an atlas of how AI is used across occupations and countries, plus a paper based on 2,600 specialized AI models and a survey of 600+ US and UK scientists (nearly half of whom report daily AI use).
  • Effort News traces Irregular, an Israeli evaluation firm, behind the recent model “rogue agent” disclosures. A Tel Aviv-based firm co-founded by Dan Lahav and Omer Nevo built the cybersecurity tests used by Anthropic, OpenAI, and Meta; their flagship investor is Dustin Moskovitz’s Good Ventures. Anthropic disclosed Irregular provided both the test harness and internet access during the July-to-September runs.
  • Philadelphia activists push back on data center siting near a former oil refinery. Amanda Silberling at TechCrunch profiles Grays Ferry, where the closed Philadelphia Energy Solutions refinery once processed up to 335,000 barrels of crude per day. Over 100 people attended a City Hall rally organized by Philly Thrive.

Worth Watching

  • Whether the OpenAI-Anthropic-Google coordination produces a formal joint statement. Bellan cites a Bloomberg-sourced angle that the effort could be formalized this week. The test is whether the standards-body framework is announced as an industry body only, with FRONTIER Act language preserved, or whether the labs go to Congress together.
  • Whether WISeR expansion to additional services clears internal CMS review. The September 8 EFF post names air ambulance, cancer treatment, MRI scans, and medications without public coverage criteria as candidates. The test is whether any of those four expansions is approved before the end of federal FY2026.
  • Whether Salesforce posts Koa weights publicly and on what license. The piece says “open-weight” but does not name a license. If a permissive license lands, Koa joins the open-source-reasoning race alongside DeepSeek, Qwen, and Llama for the late-2026 leaderboard refresh.
  • Whether iLands’ unsubscribe path actually reduces inbound volumes. Sebo’s report of re-pitches within hours suggests the fix is surface-deep. The test is whether 404 Media or another outlet re-runs the survey in 30 days and sees a drop in inbound volume.
  • Whether AI Contact Hotline filings produce visible enforcement. Mehta’s piece notes that DeepMind’s whistleblowing-agent experiment produced 24 reports and no formal consequences; the OpenAI-Hugging Face breach investigation produced “thousands” of involved agents and only 5 to 6 whistleblowers, none acting. The test is whether the hotline drives any concrete response from labs or regulators inside 60 days.
  • Whether AIUC’s first public AIUC-1 audit triggers any change at named customers. Cursor, Lovable, Harvey, and ElevenLabs are wired in. The test is whether any of them ship remediation changes inside 90 days that are traceable to a flagged finding.