Anthropic keeps Fable, rents Meta's GPUs

Anthropic's compute squeeze drove the day: it reversed plans to pull Fable 5 from subscriptions, opened $10B lease talks with rival Meta, and drew Washington's interest in who gets frontier models first. Meanwhile the open-weight camp kept applying pressure — a 2-bit DeepSeek V4 Flash matched a two-node DGX rig on one MacBook, and Britain's AI Security Institute measured the closed-model lead on cyber shrinking to months.

Anthropic backs off pulling Fable 5 from subscriptions

Starting July 20, Claude Fable 5 stays bundled in Max and Team Premium plans, but at 50% of limits that are themselves being cut 33% as the bonus-usage phase ends. Pro and Team Standard subscribers effectively lose bundled access, getting a one-time $100 credit before paying API rates. Anthropic had planned to make Fable API-only over compute-capacity concerns.

Why it matters: The reversal is a direct read on competitive pressure: GPT-5.6 Sol offers similar performance at roughly a third of the cost, and nobody pays $100-$200/month for a plan that excludes the best model. Watch whether Anthropic dials back training to free GPUs for serving.

Meta and Anthropic in talks for a $10B compute lease

Anthropic is in early talks to rent Meta data-center capacity in a deal reportedly worth ~$10B over two years, with an early-cancel option Anthropic also negotiated into its SpaceX lease ($1.25B/month for the Colossus supercomputers). The pair are LLM competitors — Meta just shipped Muse Spark 1.1, priced 75% below Claude. Anthropic would most likely take Meta's Nvidia servers rather than its custom MTIA 400 silicon.

Why it matters: Two rivals may become landlord and tenant because chip access, not ideas, is the binding constraint. For API users, more leased capacity has historically translated into higher Claude Code and API rate limits.

Trump administration wants a say in who gets frontier models first

The White House's new Gold Eagle cybersecurity initiative could act as a clearinghouse determining which organizations receive early access to OpenAI and Anthropic frontier models, per CNBC, with future rollouts potentially requiring government sign-off on partners. A White House official denied approving private releases, calling testing voluntary. The report says Claude Mythos 5 and Fable 5 were briefly blocked last month over national-security concerns before access was restored.

Why it matters: Early-access programs like Anthropic's Project Glasswing and OpenAI's Daybreak have been the labs' to run; routing them through government would reshape who can build on new models first. David Sacks warned it's 'how you lose the AI race.'

A 2-bit DeepSeek V4 Flash on one MacBook ties two DGX Sparks

In a community Terminal-Bench 2.1 run, an aggressively quantized ~80GB (2.45 bits/weight) DeepSeek-V4-Flash GGUF on a single 128GB M5 Max scored 54% versus 52% for the native FP8/FP4 checkpoint on 2x DGX Spark — a statistical tie (paired McNemar p=0.82). Separately, users report the model running with a 1M-token context on a 5090 (~650 tok/s prefill, ~17 tok/s decode), and that mainline llama.cpp b10064 now matches the old dsv4 fork, making the fork unnecessary.

Why it matters: The expensive rig mostly buys serving quality — speed, concurrency, longer usable context — not accuracy. For anyone with a big-RAM Mac, heavy quantization is far more capable than its bit count suggests.

AISI: open models now trail closed systems by four to seven months on cyber

The UK AI Security Institute's first public open-vs-closed cyber assessment finds the gap has narrowed from six-to-ten months to four-to-seven. GLM-5.2 matches February's Opus 4.6 on narrow cyber tasks; DeepSeek V4-Pro lands at Opus 4.5's level. The cost gulf is stark: a 100M-token cyber-range test ran ~$85 on Opus, ~$46 on GLM-5.2, and $1.19 on DeepSeek V4-Pro — and open safeguards were trivially bypassed by simply retrying refused tasks.

Why it matters: The window in which defenders using top closed models stay ahead of freely downloadable capability is shrinking. AISI says Kimi K3, out in late July, could close it further, albeit at higher inference cost.

Databricks hits $188B, betting on open Chinese models for coding

Databricks announced a Coatue-led round (reported ~$3B) valuing it at $188B, up from $134B just five months ago. The pitch leans on its AI reinvention: internal benchmarks across its 3,000 engineers' real tasks found GLM-5.2 now handles even the hardest coding work at lower total cost than Anthropic or OpenAI. It also found the agentic harness matters as much as the model, singling out open-source Pi for cheap context management.

Why it matters: One of the largest enterprise data vendors is publicly standardizing on open Chinese weights for production coding — and telling teams that harness choice, not just model choice, drives their bill.

First loan backed by inference chips: $400M for SambaNova silicon

AI inference cloud General Compute landed a $400M loan from Upper90, reportedly the first financing to use inference-specific chips as collateral — SambaNova's power-efficient SN50, which the startup claims runs 16x faster than GPU clouds. Upper90 pioneered GPU-backed lending with Crusoe in 2021; it's now betting the next wave is cheap inference for open models, outside Nvidia's ecosystem.

Why it matters: Capital markets are beginning to price non-Nvidia inference silicon as a financeable asset, a small crack in Nvidia's dominance and a signal that serving open models cheaply is becoming its own infrastructure category.

Browse previous days →