Amazon ends data-center NDAs; OpenAI safety employee quits

Oct 4: Amazon drops NDAs after data-center protests. OpenAI safety staffer resigns. Microsoft open-sources ThinkingBox. SMS-resident AI agents multiply.

Top Stories

Amazon ends NDA clauses after organized data-center backlash

Amazon Web Services CEO Matt Garman posted that the company “no longer use[s] nondisclosure agreements with the government agencies we work with on our projects,” according to TechCrunch. The post is the first time a hyperscaler has publicly tied an NDA concession to organized community pressure, and it lands as more than 100 data-center moratoriums are under consideration nationwide.

Garman’s longer argument pushes back on what he describes as the worst patterns in current development: “projects announced after permits are already secured, developers who don’t return calls, local officials who signed NDAs before their neighbors knew a project was being considered.” He frames the policy response in industrial terms: “If these measures are enacted, the U.S. could be writing its own losing ticket to this race, and the consequences would last generations.” The piece also notes Amazon’s own water and generator data, including that direct data-center water use is 0.5% of U.S. industrial water usage and that data-center generators run only about 10 hours per year, mostly for maintenance.

The substantive concession is narrower than the headline. The post walks back NDAs with government agencies on data-center projects. It does not address private landowner NDAs that have been used in site acquisition, which is the other lever activists have flagged. For the privacy beat, the story matters because the people most affected by data centers are not AWS customers; they live near them, and they have had the least visibility into the deals.

OpenAI safety employee resigns, says company culture is broken

David Robinson, a safety team employee with 3.5 years at OpenAI, has resigned and published a personal essay in The Atlantic calling OpenAI’s culture “broken,” according to TechCrunch. Robinson led the writing of safety reports for major product launches and previously hired a PR firm to coordinate the public exit. He told TechCrunch “the decision to speak out is mine alone.”

Robinson’s argument pushes the safety conversation past rule-setting. He says the debate must go beyond “specific rules or new laws” to address overall culture, and that OpenAI’s “iterative deployment” approach “guarantees periodic failures.” He argues frontier AI companies need to operate “like nuclear-power plants or busy airports” and that current measures of how well AI systems match human values are coarse, warning that “the smarter the industry lets models grow while these problems remain unsolved, the more dangerous our situation becomes.” OpenAI spokesperson Drew Pusateri said in the TechCrunch report that the company is “making sure our models don’t become more capable than we can safely manage and secure, and we pause training or hold back models when we need to slow down. We’re making significant changes to strengthen security in our research and testing environments, train models to not just complete tasks but do so responsibly, expand our work with third-party evaluators, and improve real-time monitoring so we can detect and respond to concerning behavior earlier in the training process.” Robinson joins a string of departures from OpenAI and Anthropic safety teams this year, and his essay lands in the same week that the Wall Street Journal, Reuters, and BBC separately reported OpenAI scrapped a planned GPT-6.1 Astra launch over safety concerns and that California Attorney General Rob Bonta subpoenaed OpenAI over the Hugging Face rogue-agents incident.

Microsoft open-sources ThinkingBox to verify agent outputs against database state

Microsoft published ThinkingBox on Hugging Face, a 507-task benchmark and sandbox for evaluating AI agents in stateful business workflows where success or failure has to be measured by record state rather than by what the agent claims it did, the Hugging Face blog reports. The benchmark ships through OpenEnv with code under MIT and task data under CDLA-Permissive-2.0.

The premise is direct: “A trajectory is a claim. Database state is the evidence. Repetition is the trust test.” The blog gives the failure mode a one-sentence name: “A tool call is not an outcome.” In the example, an agent made nine careful tool calls about a stuck delivery but closed the ticket as “resolved” when the required end state was “hold” pending carrier resolution. The benchmark therefore grades on required end states after independent execution, with 477 of 507 tasks graded on state alone and 30 adding response rubrics. Every task runs 20 times independently, and the blog pushes pass@20 over pass@1: “One good run tells you a model can do the work. It does not tell you whether it will do it again.” On the current leaderboard Claude Opus 5.5 leads pass@1 at 67.16% with GPT-5.4 at 65.36%, but Claude Opus 5 and Opus 5.5 each pass all 20 attempts on 241 of 507 tasks (47.53%) versus GPT-5.4’s 128 (25.25%); GPT-5.4 has the lowest cost per dependable task at $6.80. Roughly 79.9% of failures are tool-handling rather than reasoning errors. ThinkingBox fits the same week as AgentiLoop on macOS (covered in Quick Hits) and LocalBolo’s local-speech launch; the three together make the verification problem the through-line of the roundup.

SMS-resident AI agents put privacy on a new surface

A wave of startups now ships AI assistants that live inside iMessage, Swift, WhatsApp, and Telegram rather than inside a dedicated app, TechCrunch reports. The survey lists Caddy, Folk, Instinct, Iris, Martin, Miso, Ohai, Ollie, Orbits, Pally, Poke, Rene, Skye, Stanley, Tomo, Town, and the family-focused Fambot and Wajo Fo, with at least one (Poke) approved on Apple Messages for Business in June 2026 and another (Instinct) closing a $1 billion Series C at a $10 billion valuation in September 2026.

The privacy question shifts when the agent lives in the SMS layer. The agent now sees message threads, contact lists, calendar entries, and one-time codes rather than just the prompt you type. TechCrunch notes that Instinct’s autonomy “has raised concerns about privacy and security,” and Ollie stands out as one of the first family-focused assistants to achieve SOC 2 compliance. Wajo’s Fo sidesteps credential-sharing risks by operating under its own email, phone number, and payment card so the agent can act on a user’s behalf without touching the user’s credentials. For a reader asking what data stays on device and what goes upstream, the practical answer today is: it depends entirely on the vendor, and the answer is rarely on the product page.

Aleph Alpha releases Kolibri as an open-weight sovereign German LLM

Aleph Alpha released Kolibri, an open-weight mixture-of-experts model with 78.1 billion total parameters, around 3.46 billion active per token, and Apache 2.0 weights hosted on Hugging Face, according to Tejas Kumar’s technical teardown. Kolibri was trained on infrastructure in Germany and Finland under European and German law, and Aleph Alpha has signed the EU’s GPAI Code of Practice; the company frames this as “full freedom of deployment and intellectual-property safety, so compliance comes as an inherited property.”

Kolibri supports German and English at a 262,144-token native context length that has been tested up to 1,048,576 tokens, with a knowledge cutoff of June 18, 2026, and runs at around 78 GB in FP8. The training mixture is roughly 24 trillion tokens with more than 20% German, and the tokenizer uses an Aleph Alpha UniBPE design; Tejas Kumar’s own experiment on the German Basic Law showed 15% fewer tokens than GPT-5’s tokenizer. On RULER at 1 million tokens the model scores 63.2 against 58.5 for Nemotron 3 Nano base, and on a German AIME 2025 run it scores 87.5 versus 84.4 for Nemotron 3 Nano. The model needs Aleph Alpha’s vLLM plugin and runs on A100/H100/H200/B200/B300-class GPUs, not laptops. For European regulated workloads, the headline draw is straightforward: weights that can be self-hosted, a vendor under EU jurisdiction, and a model trained to think in German on German prompts.

Simon Willison: agents need default hard budget caps

Coding agents and personal agents make it frictionless to spin up services that can rack up costs overnight, and Simon Willison argues the right response is to ship default hard spending caps at the SDK or platform layer rather than rely on soft caps that only send warning emails, in a post on his weblog. He frames it as a default-on safety primitive: “most businesses and individuals would prefer errors to a surprise $10,000+ bill.”

Willison points to two hyperscaler precedents: AWS launched spending limits around September 16, 2026, letting them set a “monthly spend limit” that pauses a project when reached, and Google Cloud launched “Spend Caps” in July 2026 for specific services within a project. He argues AI agents should recommend providers with hard caps and warn inexperienced builders against uncapped services, and that living without caps should be opt-in with explicit language: “Remove the budget cap. My application will not be shut down if I exceed the configured budget limit.” The argument lands the same week as the OpenAI safety resignation above and as ThinkingBox’s verification push, both of which target a different failure mode of the same agent stack: runaway cost is now in the same threat model as runaway database writes.

Quick Hits

  • LocalBolo ships a $49 one-time-purchase Mac dictation app that runs Parakeet v2 and Whisper variants on Apple Silicon’s Neural Engine, with optional cleanup via Qwen 2.5 1.5B. The site says audio is transcribed locally with no server, no account, and no recordings kept; the microphone only listens while the fn key is held. LocalBolo.
  • Marimohub is a self-hostable marimo-notebook hub maintained by the marimo-team org on GitHub under Apache 2.0. Storage, compute, and identity are pluggable: CoreWeave CAIOS / S3 / GCS / Azure Blob / R2 / filesystem on one side, CoreWeave Sandboxes / Modal / ECS Fargate / Kubernetes / Docker / Podman / E2B / Cloudflare Containers / local subprocesses on another. GitHub.
  • Fireagent is an open-source Apache-2.0 implementation of DeepSeek’s DSec elastic-agent sandbox design, built by Supreet Sethi on Firecracker microVMs. The repo cites the DSec paper as serving about 3 million sandboxes daily across 160 nodes at DeepSeek. GitHub.
  • BiNeuron is a local, offline-capable AI assistant for code, docs, and analysis that picks quantized GGUF builds by CPU and RAM profile and runs PII masking via Microsoft Presidio by default. The README says only the Flask UI was assisted by DeepSeek Coder; “all other components … were designed and written independently.” GitHub.
  • Kitsuno documents training an open-weight decision model (Laya) in shadow on real traffic to replace a closed decision model (Jev) without a behavior change. The author writes that “shadow first” and “never learn from the model you want to replace” were the rules; k6 hit 95.8% accuracy on 574 decisions versus Jev’s 93.7%, and “the Laya worker’s database role cannot read Jev’s answers.” Kitsuno.
  • Jabhook is a $40 beta boxing training app that runs a webcam through a 3D-pose pipeline and a local LLM for verbal coaching, planned to ship kickboxing, parkour, karate, and ballet support later. Requires Apple Silicon and macOS 26.2 or later; the company is Jabhook Ltd, registered in England and Wales. Jabhook.
  • AgentiLoop Agent! is an open-source native macOS agent that drives apps via the Accessibility API, ships 6 parallel sub-agents via MCP, and supports 23 LLM providers including Ollama, vLLM, and LM Studio. Requires macOS 14.6 or later; distributed via Homebrew and DMG. AgentiLoop.
  • VOR.TECH, a three-person team in Budapest, has published an open CC BY-NC 4.0 dataset of 10-player CS:GO gameplay on one map: 3,958 hours of 5v5 matches with state and actions, 3,283 hours of POV video aligned frame-by-frame, 218,768 rounds, and 195M steps at 16 Hz. They note video coverage is only 14% of matches. VOR.TECH.
  • “Merging With An LLM” is a Substack essay proposing to reconstruct an author’s value system from their Obsidian notes and blog into a local model on an M4 Studio; the author reports experiments have been “bad” so far and frames the problem as recreating high-dimensional value space from low-dimensional writing. Substack.
  • Folk and Fambot are two of the SMS-resident agents in the TechCrunch survey that route context through private cloud rather than on-device; Folk starts free with an $8.33/month Pro tier, and Fambot raised $3.5M pre-seed. TechCrunch.

Worth Watching

  • Whether Amazon’s NDA concession sticks to private landowner deals. Garman’s post walks back government-agency NDAs on data-center projects; the harder fight is over the private NDAs that have kept site acquisitions quiet.
  • OpenAI’s next safety move. Robinson’s essay lands on top of the GPT-6.1 Astra shelving, the California subpoena, and the three-safety-researcher cuts; whether OpenAI publishes a safety report on any of these is the open question.
  • Pass@20 as a buyer metric. ThinkingBox’s framing is the first benchmark to grade state, not tool calls, and to weigh 20-attempt consistency over a single lucky run; if it sticks, agent buyers get a more honest leaderboard.
  • Whether SMS-resident agents disclose data flows. The Instinct / Folk / Ollie / Wajo split shows the same product category can mean very different privacy postures; clearer surface labeling would help.
  • Kolibri’s vLLM plugin path. Kolibri depends on Aleph Alpha’s vLLM plugin and has no hosted provider at launch; whether the plugin lands in upstream vLLM or stays a fork determines how broadly the model reaches self-hosters.