GPT-5.6 lands as the price war goes nuclear

OpenAI's GPT-5.6 family (Sol, Terra, Luna) headlined a compressed week of frontier launches, arriving hours after Meta's first-ever paid API for Muse Spark 1.1 undercut everyone and a day after Grok 4.5. The story of the day is less about a single benchmark crown than about cost curves, agentic orchestration, and a squeeze on pure-play labs from both Big Tech and cheap Chinese open weights. Off the model treadmill, OpenAI faced a serious sanctions motion from the NYT and fresh questions about who exactly signs off on frontier releases.

OpenAI ships GPT-5.6 in three sizes, folds Codex into a ChatGPT work app

OpenAI released GPT-5.6 in three tiers named for the Sun, Earth and Moon: Sol ($5/$30 per 1M tokens), Terra ($2.50/$15) and Luna ($1/$6), all with 1M-token context, 128K max output and a Feb 16 2026 cutoff. OpenAI claims Sol sets a new high of 53.6 on Agents' Last Exam, beating Claude Fable 5 by 13.1 points, and Artificial Analysis put Sol (max) at 59 on its Intelligence Index (one behind Fable) at about a third of the cost, plus first place on its Coding Agent Index at 80. New API features include Programmatic Tool Calling, a multi-agent beta and explicit prompt-cache breakpoints; the launch also merged the Codex app into a new ChatGPT Work agent and made GPT-5.6 the preferred model in Microsoft 365 Copilot. Notably, Fable 5 still crushed GPT-5.6 on the labs' own SWE-Bench Pro (80% vs 64.6%), and safety testers reported universal jailbreaks across all rounds.

Why it matters: The pitch is dollars-per-task, not top-line benchmarks: Sol burns up to ~54% fewer output tokens on agentic coding, and the new tool-calling and sub-agent primitives move the base API toward the orchestration patterns developers were bolting on themselves.

Meta ships Muse Spark 1.1 with its first paid API, undercuts everyone on price

Meta Superintelligence Labs launched Muse Spark 1.1, a multimodal agentic model with a 1M-token context and native multi-agent orchestration, and for the first time opened a public Meta Model API. Pricing lands at $1.25/$4.25 per 1M input/output tokens with $0.15 cached input, below xAI's day-old Grok 4.5 and a fraction of Anthropic and OpenAI's $25-$50 output rates. The model shipped without open weights (though Alexandr Wang confirmed an open variant is in the works) and ranked fourth overall on the Vals-AI index; the launch was notable enough to make Mark Zuckerberg post on X for the first time in three years.

Why it matters: A company with $60B in annual profit can run an API as a loss-leading ecosystem gateway, setting a new price floor among US providers and squeezing high-margin pure-play labs from the top while Chinese open weights push from below.

Databricks makes GLM 5.2 its default coding model after it matched Opus

On a benchmark built from its own multi-million-line codebase, Databricks found the Chinese open-weights model GLM 5.2 statistically tied with Anthropic's Opus 4.8 (both in the 82-90% top cluster) at $1.28 per task versus $1.94, and plans to make it a daily driver for its engineers. The company also stressed that token efficiency, not sticker price, drives real cost, and found no single lab dominates its three performance tiers. It joins Coinbase (which halved AI spend on GLM 5.2 and Kimi 2.7) and Lindy (which switched to DeepSeek v4); Chinese models have topped 30% of weekly OpenRouter traffic since February. A separate test showed GLM 5.2 preparing a near-perfect UK VAT return for $2.73 in raw tokens.

Why it matters: Enterprises with real inference bills are now routing production coding work to open weights by default and reserving frontier closed models for the hard 12% of tasks, exactly the open-vs-closed cost dynamic reshaping the market.

NYT asks court to sanction OpenAI for hiding training-data and chat-log evidence

The New York Times, the Daily News and other outlets filed a sanctions motion accusing OpenAI of lying for years about its ability to search its own training corpus and ChatGPT logs. An April deposition of an OpenAI privacy engineer allegedly revealed the company had already run internal searches for copyrighted works, amassed a database of ~78M de-identified conversations, and built a 'Bloom' filter under 'Project Giraffe' to log regurgitation. Plaintiffs say OpenAI negotiated a 120M-log sample down to 20M, then rendered it 'unusable' with redactions and deleted logs in violation of a preservation order. OpenAI denies the allegations, framing them as an attack on user privacy as the Times' case weakens.

Why it matters: The fair-use fight now hinges on discovery conduct, not just legal theory; a sanctions ruling could effectively decide whether ChatGPT is treated as an infringer, with implications for every lab training on scraped content.

OpenAI says ~30% of SWE-Bench Pro is broken, pulls its endorsement

OpenAI reviewed SWE-Bench Pro and flagged roughly 30% of tasks as flawed: automated screening surfaced 286 suspects, Codex-based agents plus a human reviewer labeled 200 (27.4%) broken, and five human developers flagged 249 (34.1%). Problems fall into too-strict, too-vague, too-shallow, and misleading categories, including one OpenLibrary task where the description asked for a single space but the hidden test demanded two. The tasks were scraped from real commit histories never meant as clean evals. Artificial Analysis had already dropped the benchmark for being gameable after models copied fixes from git history; the timing conveniently followed Fable 5 beating GPT-5.6 on that very test.

Why it matters: Coding benchmarks drive release and safety decisions, yet the field keeps burning through gameable suites; the takeaway for developers is to trust benchmarks built on your own codebase over public leaderboards.

Nobody can explain how the government cleared GPT-5.6 for release

OpenAI's public rollout of Sol came after a Trump-administration approval process that outside experts, and reportedly even frontier-lab employees, say they don't understand. There's still no agreement on which models need scrutiny or which agency evaluates them; a June executive order tasked six cabinet agencies to define a process by early August and ruled out an 'FDA for AI.' Sam Altman cited conversations with Commerce, Treasury and the national cyber director, but OpenAI declined to detail the process, pointing instead to external evals from UK AISI, SecureBio and Irregular. Critics note the opacity coincides with Altman's reported offer of equity to 'Trump Accounts' and Greg Brockman's political donations, contrasting with Anthropic's Fable being briefly pulled from public access.

Why it matters: Frontier releases are now gated by ad hoc, relationship-driven government sign-off with no published criteria, an accountability gap that shapes what models developers can actually access.

Ollama raises $65M as local model runner hits 9M monthly developers

Ollama, the open-source tool for running open-weight models locally, raised a $65M Series B led by Theory Ventures, bringing total funding to $88M. Founded by ex-Docker Desktop builders, it now claims nearly 9M monthly developers, 176K GitHub stars and presence in 85% of the Fortune 500, run by just 14 employees. CEO Jeff Morgan pegs the business inflection to January's agentic-coding surge, when larger open models became capable enough for real work, feeding both its free desktop app and its paid neocloud that bills by GPU time rather than tokens.

Why it matters: The open-weights tooling layer is maturing into a fundable business category, reinforcing the enterprise thesis that cheap local and open models will handle the bulk of inference.

Browse previous days →