Fable 5.1's cheaper cache, pricier tasks
Anthropic and OpenAI both shipped news on the same Tuesday: Claude Fable/Mythos 5.1 reset the coding leaderboard while quietly raising per-task costs, and OpenAI declared Astra the first model to cross its "critical" cyber threshold. Underneath the flagship launches, the day's throughline was agents eating the plumbing — Gemini's video pipeline, open-source PR queues, and enterprise cyber defense all handed to autonomous loops.
Claude Fable 5.1 cuts cache reads 75%, but tasks run ~20% dearer
Anthropic released Claude Fable 5.1 and Mythos 5.1, keeping input/output list prices at $10/$50 per million tokens while cutting cache reads from $1.00 to $0.25. Anthropic advertised savings of up to 45% on heavily agentic runs, but Artificial Analysis, a pre-release tester, found Fable 5.1 at max effort actually costs about 20% more per task than Fable 5 because it emits roughly 1.7x the output tokens. Benchmarks jumped sharply, including 52.6% on the new Terminal-Bench-Science 0.1 (up from 24.7%), and these are the first Claude models to ship with built-in watermarks plus a private-preview detection API. Fable 5.1 can now flag software vulnerabilities but still routes exploit generation to Opus; several developers reported severe rate limits and false-positive safeguard flags, and some analysts argued Fable and Mythos 5.1 are the same weights behind different safety routing.
Why it matters: The cache cut is real, but the headline savings evaporate once you count output-token bloat — read the per-task figures, not the launch post. The 'same weights, different safeguards' question also muddies which model actually earned each benchmark row.
- Anthropic's Claude Fable 5.1 promises better coding and research at up to 45 percent less (The Decoder)
- Anthropic's new Fable release is cheaper, less restrictive (TechCrunch AI)
- [AINews] Claude Fable/Mythos 5.1: new SOTA model, 75% cache price cut but 70% more output tokens (Latent Space (swyx))
- Claude Fable 5.1 made me a really nice animated pelican (Simon Willison)
- Anthropic Says New Fable AI Model Is Cheaper, Better at Coding (Bloomberg)
OpenAI says Astra is its first model to reach 'critical' cyber capability
OpenAI announced that its forthcoming Astra model crossed the Critical cybersecurity threshold in its Preparedness Framework — meaning, by its own definition, the model can independently find and exploit previously unknown vulnerabilities and chain exploits. OpenAI says it paused related training for several weeks, then resumed after adding safeguards including a 'misalignment monitor' that it concedes may occasionally flag legitimate activity. The company reports Astra scored 100% on ExploitBench and found two zero-days in a modified test, but no third party has verified these claims. A less-restricted version goes to Daybreak Blue partners such as Cisco, Cloudflare, and Palo Alto Networks at launch.
Why it matters: This is OpenAI's version of the same gated-cyber-capability playbook Anthropic ran with Mythos — and, per its own note, both a capability disclosure and a marketing claim no outsider can currently check.
- Path to Astra: critical capabilities and frontier safeguards (OpenAI)
- OpenAI Is About to Release Its First AI Model With 'Critical' Cyber Abilities (WIRED)
- OpenAI's Astra model is on the way — and very good at breaking into computer systems (TechCrunch AI)
- OpenAI to limit access to Astra's most powerful cyber tools (Axios)
Gemini's agentic video understanding cuts token use up to 88%
Google DeepMind launched agentic video understanding across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Instead of ingesting video at a fixed 1 FPS, the model decides which segments to inspect and through which modality (frames, audio, or transcript), invoking an internal tool to load only the relevant portion. Google claims up to 66% lower cost, 88% fewer tokens, and up to 7% higher accuracy on standard benchmarks, with sub-second moment retrieval and needle-in-a-haystack search over multi-hour footage. It is live via the Gemini API in AI Studio at standard token pricing — set processing to 'agentic' — with Gemini app and YouTube 'Ask YouTube' rollouts planned.
Why it matters: For anyone paying per token to search or edit long video, letting the model choose its own frames is a concrete, no-extra-fee cost lever available today.
Top AI open-source projects are closing PRs to human contributors
A Latent Space report documents projects like Flue and tldraw auto-closing external pull requests and converting them into issues or discussions, in part because so many are AI-generated. Vercel's 'software factory' of triage, fix, and review agents for the AI SDK — which had over 1,000 open issues and nearly 800 PRs — now authors 25–35% of merged PRs and closes 70–80% of issues, four weeks in. Astro's auto-triage and Ghostty's Mitchell Hashimoto predict large projects will close code contributions entirely, on the logic that if a well-specified issue can be coded by a trusted in-house agent, an external PR adds little.
Why it matters: The contribution model open source ran on for 18 years is being rewritten in real time — if you contribute to these projects, the path in is now discussion and issues, not code.
- PRs NOT Welcome: How Top AI Open Source Projects Are Managing Thousands of Contributors (Latent Space (swyx))
AfterQuery becomes YC's fastest unicorn at a $3.2B valuation
AI training-data startup AfterQuery reportedly raised a round valuing it at $3.2 billion, five months after announcing a $30 million Series A at a $300 million valuation — which YC partner Gustaf Alströmer calls the accelerator's fastest ever launch-to-unicorn. In April the company reported a $100 million annualized revenue run rate and named Nvidia, Legora, and Motif Technologies as customers. Rather than optimizing answer accuracy, AfterQuery trains models and agents to replicate how professionals like doctors and lawyers complete tasks. Forbes first reported the round.
Why it matters: The training-data layer — Mercor, Scale, now AfterQuery — keeps commanding frontier-scale valuations, a signal that expert task data, not just more compute, is the current bottleneck labs pay up for.
ChatGPT for Healthcare connects to Epic EHR records
OpenAI added an Epic integration that pulls authorized, read-only patient data — appointment notes, labs, medications — into ChatGPT for Healthcare, alongside a Healthcare Public Data plugin wiring in nine official sources including PubMed, ClinicalTrials.gov, DailyMed, and CMS Coverage. OpenAI says physicians rated 99.1% of 4,363 responses across 27 clinical use cases as safe, and more than 93% of responses per connected data source as 'good' or better on accuracy. The AI does not write back to the chart. The rollout lands amid a wrongful-death suit and a Florida pastor's near-fatal-advice claim against the company.
Why it matters: Deep EHR access is exactly the integration hospitals have been waiting for and the one liability lawyers are watching — a 99.1% safety rate still leaves a non-trivial tail on a system touching patient records.
NVIDIA and CrowdStrike build a Nemotron-based agentic cyber defense
At Fal.Con 2026, CrowdStrike and NVIDIA unveiled SafeMind, an agentic cybersecurity system pairing CrowdStrike's models and harnesses with a defensive model built on open NVIDIA Nemotron and post-trained on CrowdStrike threat data. It runs an offensive red-team agent against a defensive blue-team agent in a continuous coevolution loop on a digital twin of NVIDIA's own network. CrowdStrike claims internal evals showed its Nemotron 3 Super-based 'Blue Solano' model beat leading frontier models on accuracy at 99% lower cost. A companion product, Falcon IQ, orchestrates more than 50 agents for assessment and remediation.
Why it matters: The pitch — post-train an open model on your own security data rather than rent a closed frontier API — is a concrete argument for why defenders may prefer inspectable open weights in high-stakes domains.
Also worth a look
- BenchMIRT: What are LLM benchmarks actually measuring? (Hugging Face (Allen AI))
- The efficient frontier of LLM inference (Baseten (Hacker News))
- Paint.NET ships a clean-room Direct2D rewrite for WINE — 180,000 lines 'vibe coded' with Claude (Simon Willison)
- Codex bundles LibreOffice, Python, and Node in a 1.7GB runtime (Simon Willison)
- Python 3.15.0 release candidate 2 is here (Simon Willison)
- OpenAI's teen safety rules show Meta what comes next (Quartz)