Google's third Flash in six weeks, no Pro in sight
Google shipped Gemini 3.8 Flash — its third Flash release in six weeks — while its promised frontier Pro model stays missing, the clearest sign yet that the price-performance war is where the action is. Elsewhere the safety and policy fronts lit up: researchers sounded alarms over reporting that OpenAI's Astra uses opaque "recurrent depth" that could gut chain-of-thought monitoring, and the US DOJ filed its first brief in the AI copyright wars, siding with OpenAI on fair use. Meta's Muse Spark 1.3 and a wave of GitHub and Hugging Face engineering write-ups rounded out a dense day.
Gemini 3.8 Flash lands as Google's third Flash in six weeks — frontier Pro still MIA
Google released Gemini 3.8 Flash in two variants: a general reasoning/coding model and a defenders-only 3.8 Flash Cyber via the new Fairwind Program. Google claims 73.7% on DeepSWE v1.1 (just under Claude Opus 5's 74.0%), and Artificial Analysis scores it 59 on its Intelligence Index. Pricing holds at $0.75/$3.75 per million input/output tokens through year-end (rising to $1.50/$7.50 in January 2027), but Google concedes the model 'works harder' — Artificial Analysis clocks cost-per-task at $0.58, up ~40% from 3.7 Flash's $0.40, so per-token savings partly evaporate. Google recommends sticking with 3.7 Flash for efficiency-first work. No Gemini 3.5 Pro or Gemini 4 in sight.
Why it matters: The cheapest model at its intelligence tier is now a moving target that changes every three weeks — but 'works harder' means budgeting by task, not by token. If you optimize for spend, the old Flash may still be the better buy.
Astra's 'recurrent depth' rattles safety researchers over lost chain-of-thought
The Information reported that OpenAI's upcoming Astra uses 'recurrent depth' (a.k.a. opaque recurrence or looped transformers), cycling a query through internal layers before emitting output — leaving fewer legible reasoning traces to monitor. Redwood Research's Ryan Greenblatt called it possibly 'the single worst development for AI security/safety to date,' and Zvi Mowshowitz floated laws to head off a 'race to the bottom.' OpenAI pushed back: chief scientist Jakub Pachocki said Astra's chain of thought stays legible and its computation depth is 'within a factor of two of GPT-4,' insisting the lab remains committed to CoT monitoring. The Information adds that Anthropic and Google DeepMind are already discussing the technique.
Why it matters: Chain-of-thought monitoring is one of the few working levers for catching agent misbehavior — the same logs were central to investigating OpenAI's recent rogue-agent incident. If opaque architectures scale, that visibility shrinks industry-wide.
DOJ tells court AI training is fair use, siding with OpenAI against the NYT
The US Department of Justice filed a statement of interest in the consolidated New York Times v. OpenAI/Microsoft case — its first intervention in the AI copyright wars — arguing that training LLMs on copyrighted text is 'extraordinarily' transformative and qualifies as fair use. The brief separates training from output, calls a NYT win a threat to 'national security' and 'American prosperity,' and directly attacks the Copyright Office report that rejected blanket fair use. It carries advisory, not binding, weight. Plaintiffs (including Alden papers, book authors, and The Intercept) called it a giveaway to trillion-dollar firms; the Times notes it has spent over $30M on the litigation.
Why it matters: A bellwether case just gained the federal government as an amicus for the AI side. The ruling will shape whether every model builder needs licensing deals — and whether the data pipeline you rely on stays legal.
Meta's Muse Spark 1.3 arrives with open weights promised and a training-data discount
Meta launched Muse Spark 1.3, a model tuned for agentic and coding workloads, with open weights 'coming soon.' Per Latent Space's AI News, Artificial Analysis provisionally ranks it the #3 model in the world and puts its numbers near OpenAI and Anthropic's frontier models. Meta's pricing page lists $1.25/$4.25 per million input/output tokens for the standard tier, dropping to $0.10/$0.20 for a 'contributor' tier whose data is used to improve the products — a 90%+ discount for opting into training. Commenters on r/LocalLLaMA flag a claimed 98.1% MRCR at 512k–1M context and speculate the model may be too large to run locally.
Why it matters: If the open-weights release lands, it gives teams a non-Chinese open model at frontier-adjacent scores — and the contributor pricing is an explicit bet that developers will trade their data for a 10x cost cut.
- AINews: Muse Spark 1.3 matches GPT-5.6-Sol, confirming Meta Superintelligence as the newest Frontier Lab (Latent Space (swyx))
- Muse Spark 1.3 (Meta / Hacker News)
- Muse Spark open weights coming soon (r/LocalLLaMA)
Claude Fable 5.1's system prompt bolts the door on song lyrics and copyrighted characters
Simon Willison diffed Anthropic's newly published Fable 5.1 consumer system prompt against Fable 5. It adds a firm refusal to reproduce song lyrics, poems, or book passages 'in whole or in part' — landing days after Sony Music and Warner Chappell sued Anthropic over training on lyrics databases — plus a ban on drawing copyrighted characters or logos in any medium, including SVG and code-generated art (with a memorable 'no Sonic' example). Other changes: instructions to drop 'genuinely,' 'honestly,' and 'straightforward'; harm-reduction URLs (the first non-Anthropic links ever in a Claude prompt); and the removal of the end_conversation guidance, which Willison found still lives in an unpublished tool-specific layer.
Why it matters: System prompts are the closest thing to release notes for behavior changes, and this one reads as litigation-shaped. If you build on Claude, expect harder refusals on any lyrics or character-adjacent generation.
GitHub on cutting Copilot cost: optimize the task, not the tool call
GitHub published a detailed post on four efficiency changes to Copilot's shared agent harness, validated via offline benchmarks then online A/B tests. Key findings: naively shortening tool output (e.g. RTK) backfired because agents reran commands to recover missing context — 'we saved tokens locally and spent more globally.' Wins that stuck: dropping unused line-number prefixes from file reads (~3% lower daily inference cost per user), a meta-prompting pass that halved the task-tool prompt (~1,300 tokens/turn), selective compression of build/test noise, and batching background-task completions into results (~2.3% AI-credit savings). A separate migration cut code-review cost ~20%.
Why it matters: Concrete, measured harness engineering — the kind of numbers most vendors won't publish. The 'local metric trap' lesson generalizes to anyone building agents: token-per-call is the wrong objective.
Hugging Face reproduces 'RL over taste': training a coding model to paint watercolours
A Hugging Face engineering post openly reproduces Surya Narreddi's viral project of training an LLM to write ~150 lines of p5.brush JavaScript that paints watercolours, using TRL and OpenEnv end-to-end on HF infra. The reward is aesthetic, not verifiable: HPSv3 (a 7B human-preference model) judges whether it's a flower, and a Qwen3-VL pairwise judge scores it against a hand-rated pool of 178 paintings — 'the pool is the reward function.' Trained on Qwen3.5-35B-A3B with LoRA (all-linear, since MoE projections broke the default target modules); three reward mixes all learned. The writeup is candid about infra failures entering the reward as zeros and an OpenEnv websocket bug fixed upstream.
Why it matters: A rare fully open recipe for RLHF over subjective preference, with every artifact published. The transferable lesson: with an aesthetic reward, the bottleneck moves from hyperparameters to curating the dataset that defines 'good.'
Also worth a look
- Trump may be forced to reveal secret rules feds use for AI safety testing (Ars Technica AI)
- Decoding the new AI lingo: Loops, harnesses, squads, hill climbing (GitHub Blog)
- Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference (NVIDIA Developer)
- Agentic AI and Next-Gen Intelligence Sessions at PyTorch Conference NA 2026 (PyTorch)
- llm-gemini 0.34 adds Gemini 3.8 Flash support (Simon Willison)
- Perplexity open-sourced their Mac inference server (Lily) for Qwen 3.6 (r/LocalLLaMA)
- Qwen3.8-Flash-Next on 2x3090: 17 to 25-29 t/s with the expert-cache PR (r/LocalLLaMA)