OpenAI claims Navier-Stokes, mathematicians allege a scoop

The day belonged to OpenAI's claim that its agents cracked the Navier-Stokes Millennium Problem — a technical milestone that spent the rest of the day buried under allegations it scooped, and possibly trained on, the two mathematicians who got there first. Around it: an Anthropic researcher quit warning that both frontier labs are gambling with human survival, Meta shipped its Muse personal agent, and Chinese labs kept the open Flash-model train running.

OpenAI says 10,000 agents cracked Navier-Stokes in 88 hours; the authors it may have scooped disagree

OpenAI announced that an unreleased model it calls significantly more capable than GPT-6 Astra proved the full Navier-Stokes equations can develop a finite-time singularity, using roughly 10,000 coordinated agents over 88 hours at a cost it put 'in the millions of dollars,' with the result formalized in Lean. It says it will not claim the $1M Clay prize; the claim is unverified, and Clay's rules require peer review plus a two-year waiting period. Hours earlier, NYU's Tristan Buckmaster and Anthropic's Levent Alpoge had posted their own AI-assisted proof of a simpler forced-Euler case, and Buckmaster alleges OpenAI took up the problem only after hearing of their work, pursued the same unusual Cordoba-Martinez-Zoroa approach, and pressed him to drop Alpoge as co-author because Alpoge works at Anthropic. OpenAI denies its researchers or agents accessed the pair's data but concedes it 'cannot rule out' that de-identified data from their Codex sessions improved its models.

Why it matters: If a lab can flatten a famous open problem in days on rumor alone, possibly aided by researchers' own uploaded drafts, Terence Tao warns the incentive becomes to stop sharing promising directions at all, reversing centuries of open science and leaving mathematicians outside a few frontier labs with little left to work on.

Anthropic researcher quits the industry, and the alignment lead puts extinction odds above 10%

Jacob Coxon, a 27-year-old researcher who worked at both OpenAI and Anthropic, resigned from Anthropic and left AI entirely, writing that both labs are 'racing straight to self-improving superintelligence and gambling with our lives.' Anthropic's alignment lead Evan Hubinger publicly backed him, saying the company 'earnestly' believes AI could kill all humans and putting the odds above 10% within the decade while conceding there is no plan yet to align superintelligence. The posts landed as the Financial Times reported Anthropic withheld its latest model from the UK's AI Safety Institute.

Why it matters: These are insiders at the lab that markets itself on safety saying the quiet part out loud, even as Anthropic reportedly eyes a public listing near a $2 trillion valuation. 'We take safety seriously' and 'we're racing anyway' are being said by the same people.

Meta launches Muse, a personal agent that wants access to your inbox and wallet

Meta introduced Muse, a US-only consumer personal-AI agent that connects to a user's email, calendar, payments and other apps to book travel, fill forms, lower bills and make purchases via Stripe's Link. It runs on Meta's Muse Spark model, with each agent isolated in its own 'Secure VM,' a separate Sentinel agent mediating sensitive actions, secrets kept from the model, and a bug bounty up to $300k. Muse ships on the web, iOS, Android and WhatsApp, free with $20/month Power and $100/month Maximum tiers; Meta said day-one usage ran 10x its internal projections.

Why it matters: The pitch is that context and access, not raw model IQ, are now the bottleneck for consumer agents. But handing Meta live access to your email and payment methods is a trust ask that its FTC settlements and privacy history make harder to grant.

Chinese labs keep the open Flash-model train running: DeepSeek V4.1, Ling-VL, MiMo-X

The open-model cadence from Chinese labs did not slow. According to a translated announcement shared on r/LocalLLaMA, DeepSeek is beta-testing V4.1 Flash through its API, described as a 'new architecture' with native multimodal support and priced identically to V4 Flash; testers report roughly 2.24x faster output, though one notes the gain may partly reflect light beta load rather than architecture. InclusionAI posted Ling-3.0-flash-VL to Hugging Face, a 124B-parameter MoE with 5.5B active parameters, native image and video understanding, and a 1M-token context. And a leaked early-access email points to two more preview models, Xiaomi's MiMo-X-Pro and MiMo-X-Flash.

Why it matters: The open Flash tier, big sparse MoEs with a handful of active parameters and million-token windows, has become a near-monthly release train, and it is increasingly multimodal and agent-tuned by default. All three items here rest on community posts, so treat the numbers as claims until the weights are tested.

AWS benchmarks Blackwell G7 instances: native FP4 pays off on small MoEs

AWS published SageMaker benchmarks of its new G7 instances (NVIDIA RTX PRO 4500 Blackwell) against G5 (A10G) and G6 (L4) for 30B MoE inference. On a Qwen3-Coder-30B FP8 coding workload, ml.g7.12xlarge hit about 391 output tokens/second, 60.8% over G6 and 13% over G5, with lower P99 latency, using two GPUs and 64GB versus four GPUs and 96GB on the older families. G7 is the only generation with native FP4 tensor-core support, which AWS says gives it a structural edge on NVFP4-quantized MoE models; a separate Nemotron-3-Nano test put its cost per output token up to roughly 4.9x below G6.

Why it matters: For teams self-hosting small MoEs, Blackwell's native FP4 is starting to show up as concrete price-performance rather than spec-sheet headroom, though these are vendor numbers on the vendor's own hardware and regions.

Pathway's BDH reasons in latent space, claims ARC-AGI at $0.0007 a task

Pathway and AWS detailed BDH ('Dragon Hatchling'), a post-transformer architecture that performs reasoning inside a recurrent latent state instead of emitting chain-of-thought tokens, using brain-inspired sparse local interactions with only about 5% of neurons active at a time. Pathway says a 150M-parameter reasoning model built on it, BDH-CQ, reached 29.2% pass@2 on ARC-AGI-1 at roughly $0.0007 per task, trained on SageMaker HyperPod. The company frames the design as shifting the cost-accuracy frontier by not paying a per-token tax for reasoning.

Why it matters: Latent-space reasoning that skips the chain-of-thought token bill is one of the more concrete non-transformer bets to actually put up a benchmark number, worth watching even though the claims are the vendor's own and the model is tiny.

Browse previous days →