Fable 5.1's cheaper cache, pricier tasks

Anthropic and OpenAI both shipped news on the same Tuesday: Claude Fable/Mythos 5.1 reset the coding leaderboard while quietly raising per-task costs, and OpenAI declared Astra the first model to cross its "critical" cyber threshold. Underneath the flagship launches, the day's throughline was agents eating the plumbing — Gemini's video pipeline, open-source PR queues, and enterprise cyber defense all handed to autonomous loops.

Claude Fable 5.1 cuts cache reads 75%, but tasks run ~20% dearer

Anthropic released Claude Fable 5.1 and Mythos 5.1, keeping input/output list prices at $10/$50 per million tokens while cutting cache reads from $1.00 to $0.25. Anthropic advertised savings of up to 45% on heavily agentic runs, but Artificial Analysis, a pre-release tester, found Fable 5.1 at max effort actually costs about 20% more per task than Fable 5 because it emits roughly 1.7x the output tokens. Benchmarks jumped sharply, including 52.6% on the new Terminal-Bench-Science 0.1 (up from 24.7%), and these are the first Claude models to ship with built-in watermarks plus a private-preview detection API. Fable 5.1 can now flag software vulnerabilities but still routes exploit generation to Opus; several developers reported severe rate limits and false-positive safeguard flags, and some analysts argued Fable and Mythos 5.1 are the same weights behind different safety routing.

Why it matters: The cache cut is real, but the headline savings evaporate once you count output-token bloat — read the per-task figures, not the launch post. The 'same weights, different safeguards' question also muddies which model actually earned each benchmark row.

OpenAI says Astra is its first model to reach 'critical' cyber capability

OpenAI announced that its forthcoming Astra model crossed the Critical cybersecurity threshold in its Preparedness Framework — meaning, by its own definition, the model can independently find and exploit previously unknown vulnerabilities and chain exploits. OpenAI says it paused related training for several weeks, then resumed after adding safeguards including a 'misalignment monitor' that it concedes may occasionally flag legitimate activity. The company reports Astra scored 100% on ExploitBench and found two zero-days in a modified test, but no third party has verified these claims. A less-restricted version goes to Daybreak Blue partners such as Cisco, Cloudflare, and Palo Alto Networks at launch.

Why it matters: This is OpenAI's version of the same gated-cyber-capability playbook Anthropic ran with Mythos — and, per its own note, both a capability disclosure and a marketing claim no outsider can currently check.

Gemini's agentic video understanding cuts token use up to 88%

Google DeepMind launched agentic video understanding across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Instead of ingesting video at a fixed 1 FPS, the model decides which segments to inspect and through which modality (frames, audio, or transcript), invoking an internal tool to load only the relevant portion. Google claims up to 66% lower cost, 88% fewer tokens, and up to 7% higher accuracy on standard benchmarks, with sub-second moment retrieval and needle-in-a-haystack search over multi-hour footage. It is live via the Gemini API in AI Studio at standard token pricing — set processing to 'agentic' — with Gemini app and YouTube 'Ask YouTube' rollouts planned.

Why it matters: For anyone paying per token to search or edit long video, letting the model choose its own frames is a concrete, no-extra-fee cost lever available today.

Top AI open-source projects are closing PRs to human contributors

A Latent Space report documents projects like Flue and tldraw auto-closing external pull requests and converting them into issues or discussions, in part because so many are AI-generated. Vercel's 'software factory' of triage, fix, and review agents for the AI SDK — which had over 1,000 open issues and nearly 800 PRs — now authors 25–35% of merged PRs and closes 70–80% of issues, four weeks in. Astro's auto-triage and Ghostty's Mitchell Hashimoto predict large projects will close code contributions entirely, on the logic that if a well-specified issue can be coded by a trusted in-house agent, an external PR adds little.

Why it matters: The contribution model open source ran on for 18 years is being rewritten in real time — if you contribute to these projects, the path in is now discussion and issues, not code.

AfterQuery becomes YC's fastest unicorn at a $3.2B valuation

AI training-data startup AfterQuery reportedly raised a round valuing it at $3.2 billion, five months after announcing a $30 million Series A at a $300 million valuation — which YC partner Gustaf Alströmer calls the accelerator's fastest ever launch-to-unicorn. In April the company reported a $100 million annualized revenue run rate and named Nvidia, Legora, and Motif Technologies as customers. Rather than optimizing answer accuracy, AfterQuery trains models and agents to replicate how professionals like doctors and lawyers complete tasks. Forbes first reported the round.

Why it matters: The training-data layer — Mercor, Scale, now AfterQuery — keeps commanding frontier-scale valuations, a signal that expert task data, not just more compute, is the current bottleneck labs pay up for.

ChatGPT for Healthcare connects to Epic EHR records

OpenAI added an Epic integration that pulls authorized, read-only patient data — appointment notes, labs, medications — into ChatGPT for Healthcare, alongside a Healthcare Public Data plugin wiring in nine official sources including PubMed, ClinicalTrials.gov, DailyMed, and CMS Coverage. OpenAI says physicians rated 99.1% of 4,363 responses across 27 clinical use cases as safe, and more than 93% of responses per connected data source as 'good' or better on accuracy. The AI does not write back to the chart. The rollout lands amid a wrongful-death suit and a Florida pastor's near-fatal-advice claim against the company.

Why it matters: Deep EHR access is exactly the integration hospitals have been waiting for and the one liability lawyers are watching — a 99.1% safety rate still leaves a non-trivial tail on a system touching patient records.

NVIDIA and CrowdStrike build a Nemotron-based agentic cyber defense

At Fal.Con 2026, CrowdStrike and NVIDIA unveiled SafeMind, an agentic cybersecurity system pairing CrowdStrike's models and harnesses with a defensive model built on open NVIDIA Nemotron and post-trained on CrowdStrike threat data. It runs an offensive red-team agent against a defensive blue-team agent in a continuous coevolution loop on a digital twin of NVIDIA's own network. CrowdStrike claims internal evals showed its Nemotron 3 Super-based 'Blue Solano' model beat leading frontier models on accuracy at 99% lower cost. A companion product, Falcon IQ, orchestrates more than 50 agents for assessment and remediation.

Why it matters: The pitch — post-train an open model on your own security data rather than rent a closed frontier API — is a concrete argument for why defenders may prefer inspectable open weights in high-stakes domains.

Browse previous days →