OpenAI claims Navier-Stokes, mathematicians allege a scoop
The day belonged to OpenAI's claim that its agents cracked the Navier-Stokes Millennium Problem — a technical milestone that spent the rest of the day buried under allegations it scooped, and possibly trained on, the two mathematicians who got there first. Around it: an Anthropic researcher quit warning that both frontier labs are gambling with human survival, Meta shipped its Muse personal agent, and Chinese labs kept the open Flash-model train running.
OpenAI says 10,000 agents cracked Navier-Stokes in 88 hours; the authors it may have scooped disagree
OpenAI announced that an unreleased model it calls significantly more capable than GPT-6 Astra proved the full Navier-Stokes equations can develop a finite-time singularity, using roughly 10,000 coordinated agents over 88 hours at a cost it put 'in the millions of dollars,' with the result formalized in Lean. It says it will not claim the $1M Clay prize; the claim is unverified, and Clay's rules require peer review plus a two-year waiting period. Hours earlier, NYU's Tristan Buckmaster and Anthropic's Levent Alpoge had posted their own AI-assisted proof of a simpler forced-Euler case, and Buckmaster alleges OpenAI took up the problem only after hearing of their work, pursued the same unusual Cordoba-Martinez-Zoroa approach, and pressed him to drop Alpoge as co-author because Alpoge works at Anthropic. OpenAI denies its researchers or agents accessed the pair's data but concedes it 'cannot rule out' that de-identified data from their Codex sessions improved its models.
Why it matters: If a lab can flatten a famous open problem in days on rumor alone, possibly aided by researchers' own uploaded drafts, Terence Tao warns the incentive becomes to stop sharing promising directions at all, reversing centuries of open science and leaving mathematicians outside a few frontier labs with little left to work on.
- [AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded (Latent Space (swyx))
- OpenAI's millennium proof dispute raises the question of whether researchers can trust AI labs (The Decoder)
- What OpenAI's latest controversy tells us about the future of math (MIT Technology Review)
- OpenAI says its models solved one of math's hardest problems as researchers cry foul (France 24)
- Quoting Terence Tao (Simon Willison)
- Two dire warnings, one from Terence Tao, the other from someone who just quit Anthropic (Marcus on AI)
Anthropic researcher quits the industry, and the alignment lead puts extinction odds above 10%
Jacob Coxon, a 27-year-old researcher who worked at both OpenAI and Anthropic, resigned from Anthropic and left AI entirely, writing that both labs are 'racing straight to self-improving superintelligence and gambling with our lives.' Anthropic's alignment lead Evan Hubinger publicly backed him, saying the company 'earnestly' believes AI could kill all humans and putting the odds above 10% within the decade while conceding there is no plan yet to align superintelligence. The posts landed as the Financial Times reported Anthropic withheld its latest model from the UK's AI Safety Institute.
Why it matters: These are insiders at the lab that markets itself on safety saying the quiet part out loud, even as Anthropic reportedly eyes a public listing near a $2 trillion valuation. 'We take safety seriously' and 'we're racing anyway' are being said by the same people.
- Anthropic safety researcher says more than 10% chance AI 'could kill all humans' (BBC)
- 'Gambling with our lives': AI researcher quits Anthropic with dire warning about safety (Politico Europe)
- 'Gambling with our lives': AI researcher quits Anthropic and leaves AI entirely over what he calls a threat to humanity (Yahoo Finance)
- Anthropic Alignment Lead Warns There's '>10% Chance' AI Could 'Kill All Humans' By Next Decade (Forbes)
Meta launches Muse, a personal agent that wants access to your inbox and wallet
Meta introduced Muse, a US-only consumer personal-AI agent that connects to a user's email, calendar, payments and other apps to book travel, fill forms, lower bills and make purchases via Stripe's Link. It runs on Meta's Muse Spark model, with each agent isolated in its own 'Secure VM,' a separate Sentinel agent mediating sensitive actions, secrets kept from the model, and a bug bounty up to $300k. Muse ships on the web, iOS, Android and WhatsApp, free with $20/month Power and $100/month Maximum tiers; Meta said day-one usage ran 10x its internal projections.
Why it matters: The pitch is that context and access, not raw model IQ, are now the bottleneck for consumer agents. But handing Meta live access to your email and payment methods is a trust ask that its FTC settlements and privacy history make harder to grant.
- Meta debuts its Muse AI agent. Will consumers trust it? (TechCrunch)
- Muse - Meta's personal AI agent (Hacker News)
Chinese labs keep the open Flash-model train running: DeepSeek V4.1, Ling-VL, MiMo-X
The open-model cadence from Chinese labs did not slow. According to a translated announcement shared on r/LocalLLaMA, DeepSeek is beta-testing V4.1 Flash through its API, described as a 'new architecture' with native multimodal support and priced identically to V4 Flash; testers report roughly 2.24x faster output, though one notes the gain may partly reflect light beta load rather than architecture. InclusionAI posted Ling-3.0-flash-VL to Hugging Face, a 124B-parameter MoE with 5.5B active parameters, native image and video understanding, and a 1M-token context. And a leaked early-access email points to two more preview models, Xiaomi's MiMo-X-Pro and MiMo-X-Flash.
Why it matters: The open Flash tier, big sparse MoEs with a handful of active parameters and million-token windows, has become a near-monthly release train, and it is increasingly multimodal and agent-tuned by default. All three items here rest on community posts, so treat the numbers as claims until the weights are tested.
AWS benchmarks Blackwell G7 instances: native FP4 pays off on small MoEs
AWS published SageMaker benchmarks of its new G7 instances (NVIDIA RTX PRO 4500 Blackwell) against G5 (A10G) and G6 (L4) for 30B MoE inference. On a Qwen3-Coder-30B FP8 coding workload, ml.g7.12xlarge hit about 391 output tokens/second, 60.8% over G6 and 13% over G5, with lower P99 latency, using two GPUs and 64GB versus four GPUs and 96GB on the older families. G7 is the only generation with native FP4 tensor-core support, which AWS says gives it a structural edge on NVFP4-quantized MoE models; a separate Nemotron-3-Nano test put its cost per output token up to roughly 4.9x below G6.
Why it matters: For teams self-hosting small MoEs, Blackwell's native FP4 is starting to show up as concrete price-performance rather than spec-sheet headroom, though these are vendor numbers on the vendor's own hardware and regions.
- Benchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6 (AWS Machine Learning)
Pathway's BDH reasons in latent space, claims ARC-AGI at $0.0007 a task
Pathway and AWS detailed BDH ('Dragon Hatchling'), a post-transformer architecture that performs reasoning inside a recurrent latent state instead of emitting chain-of-thought tokens, using brain-inspired sparse local interactions with only about 5% of neurons active at a time. Pathway says a 150M-parameter reasoning model built on it, BDH-CQ, reached 29.2% pass@2 on ARC-AGI-1 at roughly $0.0007 per task, trained on SageMaker HyperPod. The company frames the design as shifting the cost-accuracy frontier by not paying a per-token tax for reasoning.
Why it matters: Latent-space reasoning that skips the chain-of-thought token bill is one of the more concrete non-transformer bets to actually put up a benchmark number, worth watching even though the claims are the vendor's own and the model is tiny.
- Pathway's brain-inspired architecture development on Amazon SageMaker HyperPod (AWS Machine Learning)
Also worth a look
- Automated agent evaluation with Amazon Bedrock AgentCore and GitHub Actions (AWS Machine Learning)
- Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic (Hugging Face)
- Qwen3-0.6B (400 MB) on a Samsung Note 8 (2017) phone drives a real desktop Chrome (r/LocalLLaMA)
- How many agents can 2x4090 actually run at once? Three weeks of llama.cpp concurrency data (r/LocalLLaMA)
- Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context (r/LocalLLaMA)
- I built Infercat: Share your local AI with friends over an encrypted p2p tunnel (r/LocalLLaMA)
- Qwen 3.8 27b with PI agent - pushed to its 3D graphic game limits (r/LocalLLaMA)