Anthropic keeps Fable, rents Meta's GPUs
Anthropic's compute squeeze drove the day: it reversed plans to pull Fable 5 from subscriptions, opened $10B lease talks with rival Meta, and drew Washington's interest in who gets frontier models first. Meanwhile the open-weight camp kept applying pressure — a 2-bit DeepSeek V4 Flash matched a two-node DGX rig on one MacBook, and Britain's AI Security Institute measured the closed-model lead on cyber shrinking to months.
Anthropic backs off pulling Fable 5 from subscriptions
Starting July 20, Claude Fable 5 stays bundled in Max and Team Premium plans, but at 50% of limits that are themselves being cut 33% as the bonus-usage phase ends. Pro and Team Standard subscribers effectively lose bundled access, getting a one-time $100 credit before paying API rates. Anthropic had planned to make Fable API-only over compute-capacity concerns.
Why it matters: The reversal is a direct read on competitive pressure: GPT-5.6 Sol offers similar performance at roughly a third of the cost, and nobody pays $100-$200/month for a plan that excludes the best model. Watch whether Anthropic dials back training to free GPUs for serving.
Meta and Anthropic in talks for a $10B compute lease
Anthropic is in early talks to rent Meta data-center capacity in a deal reportedly worth ~$10B over two years, with an early-cancel option Anthropic also negotiated into its SpaceX lease ($1.25B/month for the Colossus supercomputers). The pair are LLM competitors — Meta just shipped Muse Spark 1.1, priced 75% below Claude. Anthropic would most likely take Meta's Nvidia servers rather than its custom MTIA 400 silicon.
Why it matters: Two rivals may become landlord and tenant because chip access, not ideas, is the binding constraint. For API users, more leased capacity has historically translated into higher Claude Code and API rate limits.
- Meta in Talks to Lease Computing Power to Anthropic in Potential $10 Billion Deal (The New York Times)
- Anthropic, Meta reportedly discussing $10B data center leasing deal (SiliconANGLE)
- Anthropic in early talks with Meta to acquire compute power (CNBC)
- Zuckerberg's plan to sell excess AI compute could find its first big customer in Anthropic (The Decoder)
- Meta, Anthropic in talks for potential $10 billion compute lease deal, source says (Reuters)
Trump administration wants a say in who gets frontier models first
The White House's new Gold Eagle cybersecurity initiative could act as a clearinghouse determining which organizations receive early access to OpenAI and Anthropic frontier models, per CNBC, with future rollouts potentially requiring government sign-off on partners. A White House official denied approving private releases, calling testing voluntary. The report says Claude Mythos 5 and Fable 5 were briefly blocked last month over national-security concerns before access was restored.
Why it matters: Early-access programs like Anthropic's Project Glasswing and OpenAI's Daybreak have been the labs' to run; routing them through government would reshape who can build on new models first. David Sacks warned it's 'how you lose the AI race.'
A 2-bit DeepSeek V4 Flash on one MacBook ties two DGX Sparks
In a community Terminal-Bench 2.1 run, an aggressively quantized ~80GB (2.45 bits/weight) DeepSeek-V4-Flash GGUF on a single 128GB M5 Max scored 54% versus 52% for the native FP8/FP4 checkpoint on 2x DGX Spark — a statistical tie (paired McNemar p=0.82). Separately, users report the model running with a 1M-token context on a 5090 (~650 tok/s prefill, ~17 tok/s decode), and that mainline llama.cpp b10064 now matches the old dsv4 fork, making the fork unnecessary.
Why it matters: The expensive rig mostly buys serving quality — speed, concurrency, longer usable context — not accuracy. For anyone with a big-RAM Mac, heavy quantization is far more capable than its bit count suggests.
AISI: open models now trail closed systems by four to seven months on cyber
The UK AI Security Institute's first public open-vs-closed cyber assessment finds the gap has narrowed from six-to-ten months to four-to-seven. GLM-5.2 matches February's Opus 4.6 on narrow cyber tasks; DeepSeek V4-Pro lands at Opus 4.5's level. The cost gulf is stark: a 100M-token cyber-range test ran ~$85 on Opus, ~$46 on GLM-5.2, and $1.19 on DeepSeek V4-Pro — and open safeguards were trivially bypassed by simply retrying refused tasks.
Why it matters: The window in which defenders using top closed models stay ahead of freely downloadable capability is shrinking. AISI says Kimi K3, out in late July, could close it further, albeit at higher inference cost.
Databricks hits $188B, betting on open Chinese models for coding
Databricks announced a Coatue-led round (reported ~$3B) valuing it at $188B, up from $134B just five months ago. The pitch leans on its AI reinvention: internal benchmarks across its 3,000 engineers' real tasks found GLM-5.2 now handles even the hardest coding work at lower total cost than Anthropic or OpenAI. It also found the agentic harness matters as much as the model, singling out open-source Pi for cheap context management.
Why it matters: One of the largest enterprise data vendors is publicly standardizing on open Chinese weights for production coding — and telling teams that harness choice, not just model choice, drives their bill.
First loan backed by inference chips: $400M for SambaNova silicon
AI inference cloud General Compute landed a $400M loan from Upper90, reportedly the first financing to use inference-specific chips as collateral — SambaNova's power-efficient SN50, which the startup claims runs 16x faster than GPU clouds. Upper90 pioneered GPU-backed lending with Crusoe in 2021; it's now betting the next wave is cheap inference for open models, outside Nvidia's ecosystem.
Why it matters: Capital markets are beginning to price non-Nvidia inference silicon as a financeable asset, a small crack in Nvidia's dominance and a signal that serving open models cheaply is becoming its own infrastructure category.
Also worth a look
- The state of open source AI (Mozilla/SlashData 2026) (Hacker News)
- Google Cloud's Always-On Memory Agent replaces RAG and embeddings with continuous LLM consolidation on Gemini 3.1 Flash-Lite (MarkTechPost)
- NVIDIA Vera Rubin maximizes intelligence per dollar for post-training workloads (NVIDIA)
- Anthropic launches Claude for Teachers — and some critics are concerned (Education Week)
- Patreon stops asking AI bots not to scrape — and starts blocking them (TechCrunch AI)
- Fine-tune video and image models at scale with NVIDIA NeMo Automodel and Diffusers (Hugging Face)
- GPT-OSS-120B, Qwen 30B and Gemma 26B on an Android phone via flash streaming, CPU only (r/LocalLLaMA)
- Bonsai 27B runs locally on an iPhone — a 27B model in 3.9GB via 1-bit quantization (r/LocalLLaMA)
- Google-backed FireSat wildfire-detection satellites launch as smoke chokes US, Canada (Ars Technica AI)
- Trellis.cpp now produces high-quality image-to-3D assets, no CUDA required (r/LocalLLaMA)