Two labs cut prices an hour apart

Anthropic and OpenAI shipped new models within 90 minutes of each other, and for both the headline feature was a cheaper token, not a smarter model — the first releases since both CEOs called to "pace the frontier." Away from the launch noise, the more durable engineering story was cost: Shopify distilling a production agent to beat frontier quality at 4% of the serving bill, vLLM splitting its stack to keep older hardware alive, and Qualcomm pushing 30B-parameter models onto phones.

Opus 5.5 and GPT-6 Sol/Luna land the same afternoon, both selling on price

Anthropic released Claude Opus 5.5 at $4/$20 per million input/output tokens (down 20% from Opus 5, with cache reads 60% cheaper at $0.20), claiming Fable 5.1-level quality about 40% cheaper to run at default effort. Roughly 90 minutes later OpenAI shipped GPT-6 Sol ($2/$10) and Luna ($0.10/$0.50), Astra-derived models priced about 50% below GPT-5.6. Anthropic's benchmarks put Opus 5.5 ahead of GPT-6 Astra on Terminal-Bench 4.0 (66.4% vs 57.9%) and top of Artificial Analysis's intelligence index at 58, but at max effort it burns ~119k tokens per task versus Astra's ~27k, so the per-task saving evaporates. Artificial Analysis found GPT-6 Sol and Luna hold GPT-5.6-level intelligence with regressions on some knowledge-work evals.

Why it matters: When both frontier labs make 'cheaper tokens' the launch headline in the same hour, the competition has clearly shifted from capability to cost-per-task — and the token-usage fine print means the sticker cut isn't always a real one.

Shopify distills a production agent past frontier quality at 4% of the serving cost

In a PyTorch case study, Shopify details a continual-learning loop that mines anonymized production failures, has a panel of frontier models 'heal' them into training trajectories, and folds them back into a smaller model via supervised fine-tuning and GRPO. Its GraphQL agent, serving up to 2,000 requests per minute, is claimed to beat the frontier baseline while cutting serving cost roughly 96% — from an estimated $27M to about $1M per year. Gist-token compression shrank the static system prompt from ~6,000 to ~1,500 tokens, dropping time-to-first-token ~19% and end-to-end latency ~38% under load. The judge is calibrated against human annotators with DSPy optimizers, and inference runs on vLLM.

Why it matters: This is the most concrete published recipe yet for the 'frontier to launch, distill to scale' pattern — the same economics driving the labs' own price cuts, but done in-house against your own traffic.

vLLM forks its stack to keep older and non-NVIDIA hardware from being left behind

vLLM is introducing a separate set of 'hardware-agnostic' layers because its new 'flat' model definitions — hand-optimized for Blackwell and rack-scale systems — are becoming incompatible with fullgraph torch.compile and drop CustomOp extensibility. Frontier models like DeepSeek V4 and Kimi K3 now ship bespoke attention and kernels, and keeping them fast for out-of-tree accelerators (IBM Spyre, TPUs, AMD, Intel) or consumer GPUs was becoming a maintenance tax. The new portable layers, built in native PyTorch plus Triton/Helion, land within 3.4% of native throughput on H100 (geomean across three models) and already work through the transformers backend via USE_HW_AGNOSTIC=1.

Why it matters: As model architectures fragment, the serving layer everyone depends on is splitting into a fast path for the newest GPUs and a portable path for everyone else — a fork worth watching if you run older or non-NVIDIA hardware.

Qualcomm's Snapdragon 8 Elite Gen 6 runs a 30B MoE model on a phone

At its Snapdragon Summit, Qualcomm announced the Snapdragon 8 Elite Gen 6 and a higher-end Extreme Gen 6, both pitched around on-device AI. New sensing hubs can run models up to 200 million parameters continuously for a local voice-in/voice-out agent and speaker differentiation, while the Extreme variant can run a 30-billion-parameter mixture-of-experts model locally. Qualcomm noted the comparison to Apple's 20B MoE foundation model from WWDC. Motorola's Signature 27 will be the first device on the Extreme chip.

Why it matters: A 30B MoE running locally on a flagship phone pushes usable on-device inference well past the small-model tier, and gives app developers a real target for privacy-sensitive, offline agent features.

Open clones of Jev multiply, and Nokia ships a training-free one

The rush to reproduce TypeSafe's Jev decision models — covered here two days ago — is now a crowded field. Nokia open-sourced AnyJev, described as a training-free layer that turns any open LLM into a calibrated decision model, per a MarkTechPost writeup. A community benchmark, JevBench 1.3.0, measures 52 systems on 534 decisions and reports Jev 1.13.0 leading at 74.4, with the open SemIf (formerly OpenJev, a Qwen3.5-4B rebuild) 1.3 points behind. Simon Willison also shipped an llm-typesafe plugin, and a developer posted 'stuntd,' a local proxy that trains a small head on your own traffic to answer typed decisions in ~22ms. Most of the ecosystem evidence remains community-posted rather than independently verified.

Why it matters: For developers doing high-volume classification, routing or yes/no gating, a small local decision model can replace paid API calls — and the tooling to build one on open weights is arriving faster than the hosted product.

Browse previous days →