Anthropic goes shopping for its own silicon

A busy pre-holiday Friday: Anthropic is simultaneously courting Samsung's 2nm fabs for a custom chip and standing up its own drug-discovery programs, while Mistral open-sources a formal-verification model that saturates math benchmarks and finds real bugs. Open weights keep gaining ground — GLM 5.2 is the community's new favorite — and the UK's safety institute warns that fixed compute budgets have been systematically underselling what agents can do.

Anthropic in early talks with Samsung to build a custom AI chip

The Information reports Anthropic is exploring a custom processor built on Samsung's 2nm process and advanced packaging, and has hired Clive Chan, an early member of OpenAI's silicon team. The project is very early: no design, testing, or defined function yet, and Anthropic insists Nvidia GPUs, Google TPUs, and AWS Trainium will remain central. Samsung, SK Hynix, and Micron were strategic investors in Anthropic's $65B Series H. The move follows OpenAI's Broadcom-built 'Jalapeño' inference chip unveiled last week.

Why it matters: Every frontier lab now wants leverage over Nvidia and its own performance-per-watt story; the question is whether Anthropic can ship silicon years behind Google and Amazon without derailing its rented-compute supply lines.

Mistral open-sources Leanstral 1.5, a 6B-active prover that catches real bugs

Leanstral 1.5 is an Apache-2.0 model (119B total, 6B active) built for Lean 4 formal verification. Mistral says it hits 100% on miniF2F, solves 587/672 PutnamBench problems, and sets SOTA on FATE-H (87%) and FATE-X (34%) at roughly $4/problem versus an estimated $300+ for Seed-Prover. Beyond math, an automated Rust-to-Lean pipeline flagged 47 violated properties across 57 repos, 11 genuine bugs and 5 previously unreported, including an integer-overflow bug in the varinteger library. Weights are on Hugging Face with a free API.

Why it matters: Formal verification that runs agentically over millions of tokens and finds bugs fuzzing misses is a concrete new tool for anyone shipping correctness-critical code — and it's cheap and openly licensed.

Anthropic launches Claude Science and its own drug-discovery programs

At its 'AI for Science' event, Anthropic unveiled Claude Science, an 'AI workbench' that consolidates research tools and datasets, and said it will develop its own drugs targeting 'neglected' diseases that Big Pharma finds unprofitable. It cited demos like spotting a year-long viral contamination in minutes and flagging 32 rare-disease candidates in under an hour. Novartis's CEO framed AI as potentially cutting drug timelines from twelve years to seven or eight. Experts caution no AI-designed drug has cleared trials, and real-world experiments remain unavoidable.

Why it matters: Anthropic selling software to drugmakers while becoming a drugmaker itself is an unusual competitive posture — and a reminder that biology's slow, wet-lab bottleneck won't yield to better models alone.

GLM 5.2 crowned the new best open-weights model — if you can cool it

Community sentiment and Simon Willison's newsletter both name GLM 5.2 the top open-weights model right now. LocalLLaMA users report strong RAG and long-context reasoning, and it ranks as the best open model on niche coding/simulation benchmarks (behind GPT-5.5). One user documented a runaway 5x RTX Pro 6000 + 5090 build chasing enough VRAM to run it well, concluding it delivers but generates serious heat and will 'take over 10 years to break even.'

Why it matters: The open-weights frontier keeps closing on proprietary models, but GLM 5.2's practical footprint is a reminder that 'best open model' still means multi-GPU rigs and real thermal engineering.

UK AI Security Institute: fixed compute budgets underrate what agents can do

AISI tested frontier models across seven benchmarks at varying token budgets and found capability is a curve, not a fixed score. Raising budgets from 1M to 10M tokens lifted SWE-Bench Pro and TerminalBench success ~25%; some cyber tasks were only solved above 10M (a few above 50M) tokens. Token cost scales with human task time as a power law — a one-week task can cost billions of tokens. Newer models benefit disproportionately, steepening the estimated cyber-capability doubling rate to every 40-50 days at 50M-token budgets.

Why it matters: If your eval caps compute, you're measuring the floor, not the ceiling — and falling token prices mean capabilities that looked unaffordable get cheaper, so budget-blind benchmarks will keep surprising people.

Meta rents out excess AI compute as Zuckerberg concedes agents lag

Meta's stock jumped ~9% on plans to sell surplus AI capacity via a new 'Meta Compute' cloud business — but the move implies its $125-145B 2026 capex may exceed its needs, and rattled data-center names like CoreWeave (-13.9% in a day) and Nebius (-17%), both Meta customers. At an internal town hall, Zuckerberg admitted the agentic push 'hasn't really accelerated in the way we expected' over the past four months, while AI chief Alexandr Wang claimed an upcoming 'Watermelon' model has caught GPT-5.5.

Why it matters: The first hyperscaler to start subletting compute is a signal worth squinting at: it hints the buildout may be running ahead of demand, with extended chip-depreciation accounting propping up earnings while the party lasts.

Epoch: critical CVEs jumped 3.5x after Anthropic's Mythos vuln-discovery claim

Epoch AI reports that high- and critical-severity CVEs rose more than 3.5x in June versus the prior monthly record, following Anthropic's April announcement that its internal Claude Mythos Preview could autonomously discover and exploit software vulnerabilities. Both Anthropic and OpenAI have since launched efforts to harden critical software with frontier models before attackers weaponize them. The data is correlational, but the timing lines up with labs turning models loose on vulnerability hunting.

Why it matters: Autonomous vuln discovery cuts both ways — the same capability that patches your dependencies floods maintainers with reports, and false-positive triage becomes its own burden.

Google DeepMind buys into A24 for filmmaking-tools research

Google DeepMind and studio A24 announced a multi-project research partnership (reported at $75M, including a Google investment) to develop new filmmaking workflows and tools via A24 Labs, anchored on systems like Gemini and Veo. Coverage frames it as DeepMind borrowing A24's cultural credibility to make its AI ambitions 'feel cooler and more inevitable' — and notes a chunk of Hollywood is quietly rooting for the deal to collapse.

Why it matters: It's a bet that generative video's adoption problem is taste and trust, not just model quality — and a test of whether a prestige brand can partner with a hyperscaler without diluting itself.

Browse previous days →