Anthropic aims to out-IPO SpaceX

Anthropic's move toward a possible record-breaking public offering dominated the day, while Nvidia extended its habit of financing its own customers to Perplexity. Underneath the money, the practical threads were about inference and control: Cerebras doubled its wafer throughput, defenders got gated access to Anthropic's Mythos 5, and local tinkerers kept squeezing Qwen3.8-27B and cheap world models.

Anthropic targets a record IPO near a $2T valuation

Bankers have told investors Anthropic could raise more than $100 billion in an IPO, matching or beating SpaceX's $85.7 billion June debut, per reporting cited by The Neuron and others. That implies a valuation near $2 trillion versus the $965 billion set in its June Series H, on an annualized run rate reported around $65 billion driven largely by Claude Code. Morgan Stanley, Goldman Sachs and JPMorgan are said to lead, with an S-1 possibly filed this month and a listing as soon as October; Amazon holds roughly a 21% stake and Alphabet about 15%.

Why it matters: A debut this size would test whether the safety-first lab can out-earn the move-fast one, and the prospectus should finally quantify Anthropic's chip-supply crunch and the fallout from losing its DoD contracts.

Nvidia weighs a Perplexity stake at $30B-plus

Nvidia is in talks to invest in Perplexity at a valuation above $30 billion, more than 50% higher than a year ago, The Information reports via The Decoder. Perplexity's annualized revenue tripled from $250 million to over $750 million, credited partly to its agentic 'Perplexity Computer' and the token consumption that comes with it. The deal follows Nvidia's recent moves on Poolside, Groq ($20 billion) and Enfabrica ($900 million).

Why it matters: Nvidia is increasingly bankrolling the companies that buy its chips, a circular financing pattern worth watching as agentic products push token demand — and GPU spend — up.

Anthropic opens Mythos 5 to defenders, pledges $35M in credits

Anthropic is expanding cyber access to its Mythos-class models through partner integrations rather than direct model access: Claude Security (Enterprise public beta) now scans code with Mythos 5 and returns findings tagged by CWE, confidence and severity, with fixes applied only via Claude Code and human approval. End users never touch the model directly, receiving defined outputs like patch lists behind abuse checks. A new Defender Advantage Fund (0xDAF) puts $35 million in Claude credits toward securing open-source projects, and the Cyber Verification Program is expanding to broader dual-use work on Opus and Sonnet.

Why it matters: It's a concrete template for shipping offense-capable models without handing them over — capability delivered as narrow defensive outputs instead of raw access.

Cerebras CS-4 doubles throughput on the same WSE-3 die

Cerebras unveiled the CS-4, a rack-scale system still built on the 5nm WSE-3 chip but doubling CS-3 performance by pushing clock speed through more power and better cooling. A rack now holds three wafers instead of two and delivers up to 4,400 tokens/sec per user — claimed up to 30x faster than Nvidia GPU setups — with memory unchanged at 44GB per wafer. A modular 'Backpack' design and disaggregated inference via AMD and AWS Trainium round it out; SemiAnalysis views the networking gains as small. OpenAI already uses Cerebras for Codex Spark.

Why it matters: The gains come from brute clock scaling, not a new node — useful if you're latency-bound on agentic workloads and can actually get rack access.

Running Qwen3.8-27B: the engine, not the quant, sets your speed

A five-day macOS shootout on an M2 Max found MTPLX and llama.cpp+MTP the fastest and highest-scoring setups for agentic coding (~20 tok/s decode), with MTP speculative-decoding drafts the main lever; DFlash2/DSpark traded context for marginal gains, and vllm-mlx leaked its chain-of-thought into output. Separately, an NVFP4 build of the same model hit a 120 tok/s average with a 451K-token KV-cache on a power-limited 400W RTX 5090. Users are also reporting the dense 27B tackling niche jobs Opus 4 couldn't, like emulating a 2000s ARM point-of-sale system.

Why it matters: Same weights, wildly different wall-clock and quality depending on the serving stack — worth benchmarking your own harness before blaming the model.

A 1.57B Dreamer 4 world model, trained for $150

A hobbyist trained a playable platformer world model from scratch for about $150: 1.57B parameters, 9.6M frames, a tokenizer at 40.41 PSNR (versus Genie's reported 35.7), FVD 32.19, and roughly 144 coherent frames before drift. The key move was generating every training frame with Procgen so the true action at each step is known, fixing the weak action-conditioning that plagued an earlier Genie-based attempt. Code and site are public.

Why it matters: Interactive world models are drifting out of frontier-lab territory, and this run argues ground-truth action data matters more than scale for controllability.

OpenAI asks California to toughen SB 53, a year after backing it

OpenAI's Global Affairs team is pushing to amend SB 53, the frontier-AI transparency law it helped pass in September 2025. It wants new requirements to monitor frontier models during training and evaluation for signs they could bypass a third party's security controls or obtain confidential data, plus hardened cybersecurity across the model development process. OpenAI frames the strategy as 'reverse federalism' — states setting compatible standards that can become national policy while Congress stalls.

Why it matters: The ask centers on the exact rogue-model and cyber-exfiltration risks labs keep flagging, and keeps OpenAI in the room writing the rules it will be judged by.

The data-efficiency gap: kids learn language on a rounding error of an LLM's tokens

MIT Technology Review surveys the BabyLM effort and the 'data efficiency gap': a preteen hears roughly 100 million words, versus the 15 trillion tokens Llama 3.1 pretrained on. The 2024 champion GPT-BERT, trained on about 100 million words, still beat Llama 2 70B on one BabyLM benchmark, while popular ideas like curriculum learning underperformed and multimodal training on baby-headcam video remains weak. With easily available web text possibly running dry by the 2030s, efficiency is becoming the constraint.

Why it matters: If pretraining data is finite, learning more from less is the next real frontier — and the most plausible way universities and minority-language communities stay in the game.

Browse previous days →