OpenAI's first chip beats Blackwell
Custom silicon stole Hot Chips: OpenAI's first in-house inference ASIC posted benchmarks beating Nvidia's Blackwell, a day before Nvidia's own earnings. Away from the racks, Bill Gates declared AI has blown past every safety threshold he once thought we'd stop at, and IBM shipped a fresh batch of Apache-2.0 reasoning models while the Qwen crowd waited on Flash-Next.
OpenAI's Jalapeño inference chip beats Nvidia Blackwell in first benchmarks
At Hot Chips, OpenAI detailed Jalapeño, its first in-house accelerator, co-developed with Broadcom and built purely for LLM inference. On SemiAnalysis's InferenceX suite — verified in-lab but with numbers supplied by OpenAI — it claims 1.5x-1.9x more throughput per watt and 1.7x-3.6x lower latency than Nvidia's GB200/GB300 racks across GPT-OSS-120B, DeepSeek R1 and Kimi K2.5, all without speculative decoding. The chip taped out in November 2025, runs at 700W with HBM4, and is inference-only; it remains at engineering-sample stage with volume production not scheduled until 2027.
Why it matters: A first-generation ASIC out-performing Blackwell is unusual, and SemiAnalysis argues the fast software bring-up means 'the CUDA moat is potentially dead' — but the honest comparison is against Rubin, not Blackwell, where the two run roughly even on cost per token.
- OpenAI Jalapeño: Better Than Nvidia Blackwell (SemiAnalysis)
- OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's Blackwell and Rubin in inference benchmarks (The Decoder)
- OpenAI Jalapeno Custom AI ASIC at Hot Chips 2026 (ServeTheHome)
- OpenAI's Jalapeño chip is built for fast inference at scale, benchmarks show (TechCrunch AI)
- OpenAI Says Its New Chip Outperforms Nvidia's Blackwell As Nvidia Prepares Earnings Release (24/7 Wall St.)
Bill Gates says AI has already crossed the danger thresholds
In a new gatesnotes essay and a wide-ranging MIT Technology Review interview, Gates argues AI has passed the points at which bio, cyber, psychosocial, job-market and even loss-of-control risks were supposed to be checked. He points to models that can design novel molecules and to non-experts being able to run cyberattacks, and floats policy responses including a per-token tax, a robot tax, and 'human-reserved' jobs. He wants any model capable of making new molecules to be monitored, and the US and China to agree on it.
Why it matters: One of tech's most-quoted optimists reframing frontier models as an active security problem lands differently than the usual doomer chorus — and the token-tax idea points straight at the API bills developers pay.
- Bill Gates says we've passed AI's danger thresholds. Now what? (MIT Technology Review)
- The choices we make about AI now are critical (gatesnotes.com)
- Bill Gates was an AI optimist. Now he's scared of what could go wrong. (The Washington Post)
- Bill Gates Is Warning That A.I. Is More Dangerous Than Big Tech Will Admit (The New York Times)
Apple pitches the M5 Ultra Mac Studio as your local-inference escape hatch
Apple unveiled new Mac Studios with M5 Max and M5 Ultra (up to 512GB unified memory, 1.2TB/s bandwidth on the Ultra, a 50% jump over M3 Ultra), explicitly marketing on-device models 'without counting tokens.' A Forbes cost analysis finds the pitch shaky: a 128GB M5 Max breaks even against a $200/month subscription only in year three, and the open models that actually fit — Qwen3-Coder-Next 80B, gpt-oss-120b, Qwen3.8-27B — score well below Claude on agentic coding benchmarks, while frontier open weights like Kimi K3 (2.8T) don't fit at all. Prompt processing on Apple silicon also remains slow. Machines ship September 22.
Why it matters: The real case for buying the box is data sovereignty, not saving money — for anyone whose code or patient data legally can't leave the building, not a cheaper Claude.
IBM ships Granite 4.2 reasoning models, 3B to 30B, under Apache 2.0
IBM released Granite 4.2, its first dense, decoder-only reasoning family, in 3B, 8B and 30B sizes, each pre-trained from scratch on roughly 15T tokens with a 512K context window. Every model has a thinking/non-thinking/low-effort switch and native tool calling; the 8B and 30B additionally go through an agentic-RL stage that trains them to edit code, drive a terminal and search the web in real sandboxes. IBM also published FP8, NVFP4, MXFP4 and GGUF quantizations and vLLM/SGLang recipes.
Why it matters: A genuinely open (Apache 2.0) reasoning stack with agentic RL baked in and ready-to-serve harness configs is a rare thing at this size — usable today without a licensing lawyer.
- Granite 4.2 LLMs: How They're Built (Hugging Face)
- ibm-granite/granite-4.2-30b · Hugging Face (r/LocalLLaMA)
Qwen3.8-Flash-Next release day: a sparse MoE that might fit on a laptop
The r/LocalLLaMA community is running a release-day megathread for Qwen3.8-Flash-Next, with an estimated 15:00 UTC drop on Hugging Face and ModelScope. Per leaked and community descriptions — not yet confirmed by Alibaba — it is a multimodal MoE with roughly 176B total parameters (about 125B main weights plus 51B in n-gram embedding tables) and only ~6B active per token. Testers speculate the large, sparsely-accessed n-gram tables could be offloaded to system RAM, putting a real 4-bit quant in the 80-90GB range.
Why it matters: If the architecture holds, a 176B multimodal model that reads only a few GB per token would be unusually friendly to consumer hardware — but every spec here is a pre-release community claim until the weights land.
Nvidia pushes Groq 3 LPX into production with a contested 4x-Cerebras claim
Nvidia moved the Groq 3 LPX — the inference accelerator from the Groq team it acquired for about $20B — into full production at Hot Chips. An Artificial Analysis benchmark clocked 3,400 tokens/sec on Gemma 4 31B at 100k context, which Nvidia frames as 4x faster than Cerebras's 882 tok/s. But the comparison hides the chip count: the SRAM-heavy LPU carries just 500MB each, so the result needs at least 64 accelerators to Cerebras's one or two, uses a dense best-case model, and ignores Cerebras's newer CS-4.
Why it matters: Token-generation speed is the currency of agentic workloads, but a headline that needs 64 chips to beat someone's one or two is a marketing number, not an efficiency one.
Robot-brain startup Generalist hits $3B valuation
Generalist, founded in 2024 by former Google DeepMind researchers Pete Florence and Andy Zeng with ex-Boston Dynamics engineer Andrew Barry, reached a $3B valuation on a ~$200M extension led by 8VC, taking its Series B to $600M total. The company builds a robot-agnostic foundation model and says its new Gen 1.5 can teach robots new tasks from video demonstrations as short as 3-12 seconds. It joins a crowded field including Physical Intelligence (~$11B) and Skild AI (~$14B).
Why it matters: Investors are pricing in a 'ChatGPT moment' for general-purpose robot policies, though the data problem — no internet-scale corpus of physical actions — means that moment may still be years off.
- Robotics startup Generalist reaches $3B valuation, sources say (TechCrunch AI)
Also worth a look
- Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original (Hugging Face)
- Fully quantized NVFP4 Qwen3.8-27B with QUASAR QAD (r/LocalLLaMA)
- Meta's paid AI agent Hatch launches soon, with a new model called Watermelon due in October (The Decoder)
- Google launches Gemini for legal work to automate contracts and research (The Decoder)
- Stability AI, maker of Stable Diffusion, raises $76 million in fresh funding (TechCrunch AI)
- How loveholidays is making everyone a builder with Codex (OpenAI)
- Quoting Paul Dix on AI rewriting 1M lines of code (Simon Willison)
- AI models flub these intelligence tests. Can you fare any better? (MIT Technology Review)
- Thomson Reuters releases Thomson-1.0-Small, a law and tax focused model (r/LocalLLaMA)
- New: Llama.cpp adaptive speculation for faster inference (r/LocalLLaMA)