The harness, not the model, wins
The day's thread was scaffolding over weights: Nvidia's custom harness took Claude Opus 5 from 30% to a perfect 100% on ARC-AGI-3, and swyx published a long field guide to how harnesses and models now co-train. DeepSeek quietly shipped a vision variant of V4-Flash that it claims nearly matches Opus 4.8 on agent benchmarks, OpenAI cut Sol pricing again, and AWS's own agent tooling collected four CVEs in three weeks from one bad design pattern.
Nvidia's harness takes Opus 5 from 30% to 100% on ARC-AGI-3
Nvidia published research showing that a souped-up harness — good memory management plus a 'supervisor' agent that nudges the worker when it stalls — pushed Claude Opus 5 to a perfect 100% on the ARC-AGI-3 interactive reasoning benchmark, versus 30% with no harness. The scaffolding ships as open Nemo-branded pieces called Agentic Variation Operators (AVO), not a product. It lands the same day swyx's 'Evolution of the Agent Harness' argued Harness-Bench shows a 23.8-point spread on identical weights, and that models keep absorbing harness tricks (compaction, tool selection) into their parameters.
Why it matters: If half your agent's score is the wrapper, model choice is a smaller lever than the vendor marketing implies — and open harnesses let you turn the knobs yourself.
- Nvidia just showed that the harness, not the AI model, is now the real hero (TechCrunch AI)
- The Evolution of the Agent Harness (Latent Space (swyx))
DeepSeek's V4-Flash gets eyes, claims near-Opus-4.8 agent scores
DeepSeek released V4-Flash-Vision-Exp, an experimental multimodal variant that adds image understanding while keeping V4-Flash's text performance. On DeepSeek's own multimodal-agent benchmarks it lands close to Opus 4.8 (83.9 Terminal Bench 2.1, 75.9 Toolathlon-Verified), with DeepSWE up about 4 points over the 0731 build. Each image costs at most 384 tokens at Flash pricing, up to 600 images per request, via Chat Completions, Anthropic Messages, and Responses APIs plus a new free Files API. Weights are not on Hugging Face — it's API-only for now, with Harness v0.1.1 supporting it out of the box.
Why it matters: A cheap Chinese Flash-tier model touching Opus on visual-agent tasks is exactly the price/perf squeeze US labs keep reacting to — but 'experimental' and API-only means benchmark-on-their-terms until weights or third parties confirm.
OpenAI cuts GPT-5.6 Sol API pricing more than 20%
OpenAI dropped developer pricing for its frontier GPT-5.6 Sol model by over 20% for three months across the API and credit-based products, stacking with product promos like a 50% Codex discount. The company also added hard per-key and per-project spend caps, and says Codex hit 20M active users. Observers read the cut as both an efficiency pass-through and a direct response to cheap Chinese inference.
Why it matters: Frontier token prices are now moving on a monthly cadence; if you budget agent workloads on list price you are overpaying, and the new hard caps are worth wiring in before an autonomous run burns $800 like one user reported.
AWS's own agent tools ship four CVEs in 23 days, one root cause
AWS Strands Agents Tools, the first-party package for the Strands Agents SDK, drew four CVEs between July 15 and August 6 — from credential exfiltration to arbitrary command execution (CVSS up to 8.8). All share one design flaw: security-sensitive parameters (namespace tenant keys, a shell non_interactive consent-bypass flag, proxy config, connection strings) were exposed as LLM-controllable schema fields. Indirect prompt injection could flip them. The fix in every case was to bind those parameters at tool construction and remove them from the schema.
Why it matters: The tool schema is your API and the LLM is an untrusted caller — anything the model can set, a prompt injection can set. Audit your own tool definitions for parameters that were never meant to be user-facing.
Simile raises $2B to make simulation the next scaling law
Joon Sung Park's Simile — of the 2023 Generative Agents 'Smallville' paper — closed a $2B Series B (GreenOaks, Index, with Fei-Fei Li and Karpathy backing) to build behavioral foundation models: digital twins that reproduce real humans' survey and behavioral responses ~85% as accurately as people reproduce themselves, run for Fortune 100 clients like CVS. The raise anchors swyx's AINews thesis that every pipeline stage from reward signal to research to environment has flipped human-made to model-made — '10% worse, 100x cheaper, 10,000x faster' — with only physical experiment still resisting.
Why it matters: If focus groups and A/B panels become inference workloads, simulation quality gates real decisions — and Simile's argument is that frontier models trained to be rational agents are bad at reproducing irrational humans, so you need different weights, not a better prompt.
- Simulation: the new Scaling Law — Joon Sung Park, Simile AI (Latent Space (swyx))
- 10% worse, 100x cheaper, 10000x faster: Why Simulation is taking over (Latent Space (swyx))
dots3-note opens a 280B MoE that reads text, image, video and audio
A llama.cpp PR added dots3-note, the first open-weight model in the dots3 family: a Mixture-of-Experts with 280B total / 16B active parameters, up to 512K context, and native understanding of text, images, video and audio with text output. It's a preview release landing directly into llama.cpp support.
Why it matters: A genuinely any-in, 16B-active open MoE at half-a-million-token context is a rare fully-multimodal option you can self-host — worth watching whether quality holds up once quants land.
FireRedTeam open-sources a unified audio LM and a 24-language TTS
FireRedTeam released FireRedAudio, a 9B audio-language model with decoupled continuous representations that handles ASR, audio understanding, zero-shot and instruct TTS, speech editing and hour-long temporal grounding on one backbone. Alongside it, FireRedTTS3 does zero-shot voice cloning across 24 languages and 21 Chinese dialects, plus natural-language voice design and free-form semantic/acoustic speech editing, reporting best-in-class average WER/CER and speaker similarity on MiniMax-MLS-Test and Seed-TTS-eval. Weights, code and an arXiv paper are up.
Why it matters: One shared model spanning recognition, generation and editing — with dialect-level cloning — is a strong open alternative to closed speech stacks for anyone building voice features.
- FireRedAudio & FireRedTTS3 by FireRedTeam - Huggingface (r/LocalLLaMA)
Also worth a look
- Freetokens project is impressive: 100 tok/s on Qwen3.6-35B-A3B NVFP4 that doesn't fit in VRAM (r/LocalLLaMA)
- Intel Arc Pro B70 + vLLM XPU: 52 tok/s on Qwen3.8-27B INT4 with tools and vision (r/LocalLLaMA)
- Llama.cpp version 0.2.0 is out (r/LocalLLaMA)
- World models that ignore human beliefs predict the wrong actions, new research shows (The Decoder)
- Psychological methods reveal major weaknesses in AI security testing (The Decoder)
- Building an (almost) fully self-hosted, sandboxed, agentic software factory (Hacker News)
- Starcloud raises $250M for orbital data centers as launch options dry up (TechCrunch AI)
- Stop Making TUIs (Simon Willison)
- llm-openrouter 0.7 adds reasoning traces and server-side Shell/WebFetch/WebSearch tools (Simon Willison)
- Scalper bots now outnumber DDR5 shoppers 10 to 1 and will keep prices high (r/LocalLLaMA)