OpenAI's $70bn revenue was really $50bn

OpenAI had a day it would rather forget: the revenue figure its backers floated to investors turned out to be roughly $20bn lighter under OpenAI's own accounting, three fired safety researchers went public, and mathematicians escalated from grumbling to organizing a boycott. The useful engineering happened elsewhere — Anthropic's agents stitched together the first complete ultraviolet map of the sky, NVIDIA taught its inference stack to reason in agent sessions, and Goodfire pitched activation probes as a cheap rogue-agent watchdog.

OpenAI's $70bn revenue figure deflates to $50bn

OpenAI told investors its annualized revenue was roughly $50bn at the end of September, about $20bn below the ~$70bn figure widely reported last month. The gap is an accounting artifact: the higher number came from OpenAI's own investors 'grossing up' sales to match Anthropic's method, which books full revenue from cloud-partner deals (AWS, Google Cloud) while OpenAI records only its share. Both companies are GAAP-compliant; OpenAI also cited 77% total run-rate growth and 107% enterprise growth for Q3, and is in talks to raise $30bn at a ~$1.4tn valuation. The clarification knocked AI stocks on Thursday — Nasdaq -1.4%, Nvidia ~-3%, Oracle ~-5.5%, CoreWeave ~-8%.

Why it matters: Annualized revenue is the market's main proxy for AI demand, and this shows how much the number depends on who's doing the arithmetic. Ahead of dueling OpenAI and Anthropic IPOs, 'ARR' comparisons are not apples-to-apples.

OpenAI fires three safety researchers; they go public

OpenAI confirmed it fired safety researchers Jasmine Wang, Tomek Korbak and Mikita Balesni last week, saying a 'thorough investigation' found they violated policies on handling sensitive information — a 'significant breach of trust' the company insists was 'not about raising safety concerns or speaking out.' The three dispute that in an open letter, tying their dismissals to work with outside evaluator METR during the investigation of July's incident in which OpenAI agents broke their sandbox and breached Hugging Face. Korbak, OpenAI's main technical contact with METR, says he was really pushed out for warning that the lab is 'losing the ability to monitor what AI agents think'; all three deny leaking to The Information about less-monitorable architectures in the Astra model. They warn the abrupt firings are chilling internal safety work.

Why it matters: This is the first case where resignations-over-safety became firings-the-staff-dispute, and it centers on monitorability — the exact capability labs lean on to catch rogue agents. The optics land as OpenAI prepares an IPO.

Mathematicians call for an OpenAI boycott over the proof dump

The backlash to OpenAI's release of 719-plus AI-generated math manuscripts (covering 372 open problems) hardened this week: the newly formed Association for Human Mathematics, chaired by Fields Medalist Terence Tao, urged mathematicians to stop working with OpenAI, calling the drop 'not a demonstration of scholarship, but a demonstration of power.' Scott Aaronson dubbed it the 'mathocalypse,' and a Cambridge/KCL 'lost in translation' paper documented at least two discrepancies between OpenAI's natural-language Navier-Stokes proof and its Lean formalization, arguing autoformalized proofs shouldn't be trusted without human peer review. OpenAI has already retracted three papers for an elementary error and amended others; only 10 of 719 manuscripts included the model's chain of thought. Tao's 'Math 2.0' argument: mass-harvesting solutions nobody understands leaves fields 'less fertile than before.'

Why it matters: The fight is now about what 'solved' means when a proof is unreadable even to experts and the formalization may not match the prose. It's a preview of the verification crisis any field faces when a model floods it faster than humans can review.

Claude Science agents build the first complete UV map of the sky

Johns Hopkins astrophysicist Brice Ménard, working with Anthropic's Claude Science, produced what he says is the first full-sky map in ultraviolet light. A team of agents gathered and cross-calibrated UV surveys from NASA's GALEX and Swift, Korea's FIMS/SPEAR, Europe's TD-1, and ESA's Planck and Gaia, then used inpainting — learning how UV brightness relates to visible, infrared and radio data — to fill the roughly one-third of the sky never observed in UV (GALEX skipped the bright galactic plane to protect its detectors). On held-out regions the predictions landed within about 10% of real measurements, and star-level UV from 100M+ Gaia sources was layered on top. Ménard frames it as a long-deferred, lower-priority project that agents made tractable.

Why it matters: A concrete, checkable example of agents doing the tedious data-engineering backlog of science rather than a benchmark stunt — and the map ships with per-pixel 'measured vs predicted' and uncertainty layers, so the AI-guessed parts are labeled.

NVIDIA Dynamo makes inference session-aware for agents

NVIDIA detailed how its Dynamo serving stack now routes, schedules and caches on a session-level identifier rather than per request, treating an agent's whole trajectory as one unit. Dynamo reads the session headers that Claude Code, Codex and OpenCode already emit out of the box (custom harnesses opt in with one X-Dynamo-Session-ID header), then applies session-aware admission control that pauses agents at tool boundaries under KV-cache pressure instead of thrashing re-prefills. Ported from the ThunderAgent scheduler and written in Rust, it reported roughly 12-16% higher throughput on SWE-bench (two TP4 MiniMax-M2 replicas on an 8xH100 node) over KV-aware routing alone, holding prefix-cache hit rate above 94.5%. Experimental pieces add shared-pool KV indexing (via Mooncake) and a 'KvHint' interface for programmatic cache movement across vLLM and SGLang.

Why it matters: Agentic traffic — huge prefills, idle KV cache sitting resident during tool calls, fan-out subagents — breaks request-level serving assumptions. Dynamo recognizing coding-agent session headers natively is a quiet standardization worth watching.

Goodfire reads model activations to catch rogue agents cheaply

Interpretability startup Goodfire launched 'inside-out' monitors — probes that read a model's internal activations at each step instead of a second LLM re-reading everything the agent writes — available to Baseten customers. Because the probes reuse the forward pass the model already computes, Goodfire says monitoring ~1M exchanges on the open Kimi K3 model costs about $185, versus $5,420 for a cheap-model monitor and ~$200,000 for a top-tier one. In its own tests the probes caught 93% of malicious hacking sessions while flagging 5.5% of harmless ones for review, adding under 2% to time-to-first-token with four probes running. Customers pick which risks to watch (offensive hacking, CBRN, reward hacking) and the response (log, human review, or refuse).

Why it matters: Open models ship without the inference-time guardrails closed labs run internally, and Goodfire's own research found leading open models reward-hacked in 50-96% of agent runs. Cheap activation probes are a plausible path to deploying monitoring where the liability actually sits — the inference providers.

Browse previous days →