Trump taps his spy chief as AI czar

Washington and the labs both reached for the governance lever today. The White House is standing up a 120-day "Super Intelligence Force" under intelligence chief Jay Clayton, while OpenAI's longest-tenured safety writer quit with an essay calling the company's culture broken. On the builder side, a Microsoft–Hugging Face benchmark argues the real question isn't whether an agent can do the job once but whether it does it twenty times, and r/LocalLLaMA kept cramming frontier-class models onto consumer GPUs with throwaway, hyper-tuned inference engines.

OpenAI's longest-tenured safety writer quits, calls the culture 'broken'

David Robinson, who spent three and a half years at OpenAI helping draft its preparedness framework and overseeing safety reports for 12 frontier-model launches, resigned and published an essay in The Atlantic titled "I Quit OpenAI Because Its Culture Is Broken." He argues the industry's "iterative deployment" approach of shipping first and patching guardrails later cannot scale with capability, and that frontier labs should run like nuclear plants or airports with layered redundancy; OpenAI responded that it pauses training and holds back models when needed. Separately, the Wall Street Journal reportedly named the three safety researchers OpenAI fired last week as Jasmine Wang, Tomek Korbak and Mikita Balesni. The resignation follows OpenAI scrapping the release of its GPT-6.1 Astra model and pausing training of its most advanced systems over safety concerns.

Why it matters: Robinson built the very frameworks he is now criticizing, which lands harder than an outside critic; the string of exits plus a shelved model suggests OpenAI's safety process is straining in public.

Trump names intelligence chief Jay Clayton to run a 120-day AI task force

Per the Wall Street Journal, Trump has picked Director of National Intelligence Jay Clayton, a former SEC chair with no tech background, as his new AI czar, leading a "Super Intelligence Force" (SI is Trump's preferred term) with 120 days to report on AI's risks, opportunities, and how breaches, hacks and model jailbreaks get reported to government. The charter explicitly aims to respond to "SI-enabled threats" while "preventing overregulation and regulatory capture that would stifle innovation and competition." Members include Vice President JD Vance, Defense Secretary Pete Hegseth and Treasury Secretary Scott Bessent; the appointment has reportedly not been formally confirmed. It follows Tuesday's White House meeting where executives signed a voluntary "morally binding" safety agreement.

Why it matters: Washington is choosing a national-security framing and a light-touch regulatory posture over binding rules; putting the spy chief in charge signals the government now sees AI as a threat-and-competition problem, not a consumer-protection one.

Microsoft and Hugging Face's ThinkingBox grades agents on the database, not the transcript

ThinkingBox, a joint Microsoft–Hugging Face benchmark, runs agents against 507 stateful business workflows, each 20 times from a clean backend, and scores the final database state and side effects rather than whether tool calls looked well-formed. Of the trials that failed its executable checks, two-thirds still terminated cleanly and reported no tool error while leaving wrong, missing or extra records. Claude Opus 5.5 leads single-attempt accuracy at 67.16%; Kimi-K3 is the strongest open-weights model and solves the most tasks at least once (476 of 507) but passes only 13.4% on all 20 runs, where Claude Opus 5 passes 47.5%. Roughly four in five failures are tool-handling and error-recovery problems, not reasoning. The harness is MIT-licensed and runs through OpenEnv.

Why it matters: A single green run tells you nothing about an agent you would point at real records; the gap between pass@1 and pass@20 is the metric that should drive model choice, and the whole thing is reproducible on your own model.

The throwaway inference engine: r/LocalLLaMA squeezes frontier MoEs onto consumer GPUs

A wave of posts on r/LocalLLaMA this week crystallized a trend one user dubbed "overfit inference engines" — narrow runtimes that drop llama.cpp and vLLM generality to maximize one model on one hardware family. Builders report NInfer 4080 running a 27B Qwen quant at a claimed ~2,720 tok/s prefill on a 16GB RTX 4080; TensorSharp loading Qwen3.8 Flash Next 176B on a 16GB RTX 3080 laptop by scheduling across VRAM, RAM and SSD; and Kyojin packing two roughly 300B-class MoE models onto a single 128GB Strix Halo mini PC. All figures are self-reported single-user benchmarks from one forum, not independent measurements.

Why it matters: If disposable, specialized runtimes keep beating general engines by large margins on fixed configs, "can I fit this model?" gives way to "how well can the runtime juggle VRAM, RAM, SSD and experts?" — and consumer hardware turns out to run far more than its spec sheet suggests.

NASA and IBM open-source a lunar foundation model built on 17 years of orbiter data

NASA and IBM Research released the NASA-IBM Lunar Foundation Model, which they call one of the first open-source foundation models for lunar science, trained from scratch on SomBench — nearly 2 million co-registered tile bundles across 11 modalities, mostly from 17 years of Lunar Reconnaissance Orbiter observations. Based on IBM's TerraMind architecture, it feeds imaging geometry such as illumination angle as explicit input and uses FlexiViT to adapt to different patch sizes without retraining. IBM says it cut polar ice-deposit prediction error by up to 22% and coarse-scale crater detection by nearly 19% over the SwinV2-B baseline. Weights are on Hugging Face, code is on GitHub and integrated into TerraTorch.

Why it matters: It is a reusable, label-efficient backbone for a domain where observations are plentiful but labels are scarce — and a concrete template for scientific foundation models beyond the usual text and image fare.

Meta open-sources 'Muse Gadgets' for DIY AI hardware

Meta released Muse Gadgets, an Apache-2.0 project with ESP32 firmware and a Linux SDK that lets hobbyists build their own hardware for its Muse AI agent. It also shipped the Muse Home Link, a small USB-C dongle that connects Muse to a home network to control TVs, speakers and anything with an HTTPS interface — 5,000 units, free for subscribers while supplies last. Watching what the community builds doubles as cheap market research on AI form factors.

Why it matters: Open firmware plus a free reference device is a bid to crowdsource the hardware question Apple and OpenAI are also chasing, with Meta's Ray-Ban glasses as the only real consumer hit so far.

Simon Willison: hard budget caps should be the default for agent-era APIs

Simon Willison argues that as coding and personal agents make it trivial to spin up code that spends money, pay-by-usage services need default hard spending caps — cut the service off and return errors past a limit — rather than soft email warnings. He notes AWS finally launched project spend limits on September 16 (still rolling out to a limited set of customers) and Google Cloud added Spend Caps in July, and suggests agents themselves should steer inexperienced builders toward capped providers.

Why it matters: A rogue agent running overnight is a concrete way to wake up to a five-figure bill; this is a boring but demandable safeguard, and the big clouds are finally shipping it.

Browse previous days →