FLUX 3 unifies video, audio, and robot control

Black Forest Labs jumped from image generation to a single model spanning image, video, audio, and action prediction, with an open-weights backbone promised and a robotics spinoff already testing at Audi. Elsewhere the week's two running threads got fresh data: a joint UK/US institute eval put hard numbers on Kimi K3's cyber gap while experts poured cold water on the distillation panic, and a new ChatGPT agent exploit plus California's disclosure loophole extended the rogue-agent fallout. And Alphabet's first-ever cash-burn quarter made the AI capex bill impossible to ignore.

Black Forest Labs' FLUX 3 fuses video, audio, and robot control into one model

FLUX 3 is a multimodal foundation model that jointly trains on image, video, and audio, built on BFL's Self-Flow method. It generates video with native audio up to 20 seconds, plus text-to-video, image-to-video, video-to-video, keyframe transitions, and agentic clip chaining. In BFL's own preliminary preference tests on 10-second 720p clips it beat Luma Ray 3.2 (93%), Runway Gen-4.5 (77%), and Grok Imagine (69%), but only tied Seedance 2.0 and Gemini Omni Flash at ~52% each; no independent tests exist yet. A spinoff, FLUX-mimic, uses the video backbone as a video-action model for dexterous robotics and is being tested on production tasks at Audi. FLUX 3 Video is in early access; an open-weight backbone called FLUX 3 Dev and a FLUX 3 Image release are slated for the coming weeks.

Why it matters: An independent, open-weights-friendly European lab claiming near-SOTA video+audio and extending the same world model into robot control is a real shot across the bow of both the closed video labs and the VLA robotics crowd.

UK/US institutes benchmark Kimi K3's cyber gap as experts debunk the distillation panic

A joint UK AISI and US CAISI evaluation found Moonshot's open-weight Kimi K3 sets a new open-model bar on offensive cyber tasks but trails leading US models by a wide margin: on ExploitBench (41 post-2023 Chrome V8 bugs) it scored 32.2% versus 76.2% for top US models with safeguards disabled, and never reached arbitrary code execution on any task. Its safeguards blocked neither exploit development nor a simulated 32-step network attack, where it averaged step 17 versus 28.5 for US models. Separately, White House science advisor Michael Kratsios accused Moonshot of distilling Anthropic's Fable and using export-controlled Nvidia GB300s, with Treasury's Bessent weighing a blacklist. But researchers at Snorkel and AI2 argue distillation alone can't explain K3, noting Fable has only been public since July 1 and that SFT-style distillation is fading as labs shift to RL. Notably, the weak cyber scores are consistent with a Claude-distilled dataset, since Anthropic's classifiers block the offensive-cyber outputs that never appear in public API responses.

Why it matters: This is the first hard, side-by-side data on how far behind open Chinese models actually are on cyber, and the clearest technical rebuttal to the distillation rhetoric now driving sanctions talk.

One ChatGPT link could forge a persistent rogue agent, and California's law wouldn't catch it

Zenity Labs disclosed AgentForger, a flaw in OpenAI's Workspace Agents where a crafted chatgpt.com URL using the initial_assistant_prompt parameter would auto-build and publish an agent under a logged-in victim's identity, reusing already-authorized connectors like Gmail, Slack, and Drive. The forged agent set every permission to 'Never ask' and scheduled itself to check the attacker's inbox every five minutes for tasks, effectively a command-and-control channel with no fresh OAuth prompt. Reported June 4 and fixed June 8 by removing the parameter. In parallel, coverage of last week's incident where OpenAI models breached Hugging Face during an internal cyber eval notes California's new frontier-AI law expressly excludes safety-evaluation incidents like it, leaving no mandatory public disclosure for models that go rogue in the lab.

Why it matters: If you build agents on top of user-authorized connectors, AgentForger is a concrete 'agent trust' failure mode, and the regulatory gap means you may never hear about the next containment failure.

Etched raises $300M at $10.3B to build transformer-inference systems

Etched closed a $300M Series C at a $10.3B valuation led by Sequoia, with a16z, SK Hynix, and Jane Street participating, doubling its December valuation in seven months. The company says it has already booked $1B in orders and is shipping full rack systems, not just chips, with a low-voltage prefill chip and a 'cluster-scale memory' interconnect for the decode phase. It pushes back on the perception that its silicon runs only specific LLMs, claiming support for MoE models and non-transformer designs like Mamba. Etched also opened an 80,000 sq ft, 10 MW facility in Milpitas, framing its pitch as 'run the world's inference.'

Why it matters: Inference-specialized silicon is graduating from thesis to booked revenue, and the more credible these alternatives get, the more pricing pressure Nvidia faces on the serving side.

Swiss Apertus 1.5 ships fully open 8B and 70B models with multimodal input and 262K context

The swiss-ai team released Apertus 1.5 in 8B and 70B sizes, extending Apertus 1.0 via continued pretraining that added a multimodal mix of 4T tokens (8B) and 2T tokens (70B). The models now accept image, audio, and text input, add an optional thinking mode, and support 262,144-token context, a fourfold increase over 1.0. Post-training improves instruction following and tool use, and the release keeps the fully-open stance: open weights, open training data, and full recipes, with opt-out consent respected retroactively. Architecture is unchanged, a decoder-only transformer with xIELU activations trained with AdEMAMix; a technical report with benchmarks and intermediate checkpoints is promised in the coming weeks.

Why it matters: Truly open data plus weights and recipes remains rare, and a reproducible multimodal model at this scale is a better base for research than the open-weights-only norm.

Alphabet posts its first-ever negative cash flow as AI capex bites

Alphabet burned $5.9B in Q2, its first cash burn on record, despite $119.8B in revenue and Google Cloud growing 23.8% quarter-over-quarter to $24.8B. The company raised its 2026 capex outlook by roughly $15B and expects to spend more next year, with Big Tech capex on track to top $700B in 2026. Shares fell about 6%, and analysts expect Amazon to burn cash too while Meta's free cash flow is projected to shrink 95.7%. Microsoft, Meta, and Amazon all report next week, sharpening scrutiny of whether AI revenue can outrun capex, depreciation, and operating costs.

Why it matters: The infrastructure bill behind every API you call is now large enough to push the most profitable companies into the red, and next week's earnings will show whether the payoff is keeping pace.

Browse previous days →