Hardware & chips
71 stories on this topic, newest first.
NVIDIA and Microsoft rebuild the Windows PC around local models and background agents
At a San Francisco event, Microsoft and NVIDIA detailed RTX Spark-powered Windows machines: the Surface Laptop Ultra starts at $2,599 with up to 128GB unified memory and up to 1 petaflop of FP4 compute, laptop preorders open now and shipping Oct 16. RTX Spark pairs a Blackwell RTX GPU (up to 6,144 cores) with a Grace CPU at 600 GB/s, pitched to run models like Qwen 3.8 Flash Next on-device. Microsoft also shipped Execution Containers (MXC), OS-level sandboxing so agents can run persistently in the background, and previewed a GB300-based DGX Station for Windows with 748GB coherent memory and up to 20 petaflops FP4. Dell, HP, Lenovo, Acer, ASUS, MSI and Gigabyte have systems coming.
Why it matters: This is the first serious attempt to make Windows a first-class target for local inference and always-on agents rather than a cloud terminal — and Microsoft is dangling $1,000 MacBook trade-ins to pull developers off Macs.
OpenTPU: an AI-designed open-source accelerator that runs its own inference
A Show HN project, openTPU, puts a full AI accelerator in one readable monorepo — SystemVerilog RTL, a custom ISA, a bit-exact simulator, a kernel language and compiler, and host software driving a real PCIe FPGA card (Xilinx Kintex-7). The design runs ten modern models with real weights on a ~$ scavenged Inspur card, producing tokens bit-for-bit identical to the simulator, and streams MoE experts from host storage for models larger than the card's 4GB. Decode is DRAM-bound at 82-85% of DDR3 peak; the authors frame it as both a research artifact and a teaching tool.
Why it matters: It's a rare end-to-end, auditable look at how an accelerator actually works, from a Python matmul down to the wires — and a concrete data point on how far AI agents can get at hardware design. Good reading for anyone curious about inference bottlenecks beyond the GPU.
- OpenTPU – An open-source AI accelerator, developed by AI (Hacker News)
Builders run Qwen3.5 on ~$300 FPGA boards scavenged from dead crypto miners
According to Startup Fortune and the maintainer's own write-up, a developer going by Nero7991 built a transformer inference engine in raw VHDL (llm.vhdl, MIT) that runs Qwen3.5-9B at INT4 on the SQRL FK33, an ex-mining Xilinx FPGA card with 8GB of HBM2 now selling around $280-300 on eBay. At 75MHz a two-card pipeline reportedly manages about 2.5 tokens/sec on the 9B, with Qwen3.8-27B as the target across two larger dies; a related project, coreyhahn's fable5_llm on an older BCU-1525, is said to hit 8.18 tok/s. Nobody is pretending these beat a modern GPU — the point is a slow, DIY path around GPU scarcity.
Why it matters: It's a vivid demonstration that the 'fixed' GPU-supply constraint has ugly but real workarounds if you'll write your own instruction set and accept single-digit tokens per second.
Meta open-sources 'Muse Gadgets' for DIY AI hardware
Meta released Muse Gadgets, an Apache-2.0 project with ESP32 firmware and a Linux SDK that lets hobbyists build their own hardware for its Muse AI agent. It also shipped the Muse Home Link, a small USB-C dongle that connects Muse to a home network to control TVs, speakers and anything with an HTTPS interface — 5,000 units, free for subscribers while supplies last. Watching what the community builds doubles as cheap market research on AI form factors.
Why it matters: Open firmware plus a free reference device is a bid to crowdsource the hardware question Apple and OpenAI are also chasing, with Meta's Ray-Ban glasses as the only real consumer hit so far.
A one-person vLLM fork gets Qwen running on Huawei's Ascend cards
In a detailed build log on r/LocalLLaMA, developer /u/matteiuspi reports taking two passively-cooled Huawei Atlas 300I Duo cards (96GB each, enumerating as four 48GB Ascend 310P devices) from incoherent ~1 tok/s output to roughly 30 tok/s single-stream and about 61 tok/s aggregate at four-way concurrency on Qwen3.8 Flash-Next, via his own forks of vLLM and vLLM-Ascend. He says the W4-packed Ascend service matched an RTX 6000 Pro llama.cpp reference at 140/198 (70.71%) on GPQA Diamond, though the Nvidia card was far faster per request. He also claims Claude repeatedly refused to help because the hardware is Chinese.
Why it matters: Non-CUDA inference is still mostly bring-it-yourself kernel work by lone developers. These are one person's unverified benchmarks, but they suggest MoE models are the sweet spot for cheap, high-memory accelerators.
OpenAI and Synopsys build GPT-Synopsys to drive EDA chip-design tools
OpenAI and EDA vendor Synopsys signed a multi-year partnership to co-develop GPT-Synopsys, a specialized model trained to operate Synopsys' electronic design automation tools directly, reasoning about chip design and verification and iterating toward power/performance/area targets for engineer review. The model runs on OpenAI infrastructure, the deal includes revenue sharing and joint go-to-market, and early engagements with semiconductor customers are underway. OpenAI says customer design data won't be used for training.
Why it matters: This pushes agents from calling EDA tools to being expert users of them, and pairs with OpenAI's Broadcom and Jalapeno chip work; the lab wants better silicon to run its own models, and chip-design flows are a high-value, closed enterprise market.
DeepSeek open-sources a TileLang toolkit to chip away at CUDA on Huawei Ascend
DeepSeek released open-source programming tools for Huawei's Ascend chips, centered on TileLang, a language it pitches as simpler to program than Nvidia's CUDA while still extracting full hardware performance. The release, which Huawei 'fully supported,' includes compute and inter-chip data libraries and optimizes a 128-chip Ascend 950 supernode; TileLang is now DeepSeek's main tool for its AGI work. SemiAnalysis has called CUDA's moat 'potentially dead' after OpenAI's Jalapeno inference chip, but still finds Nvidia ahead on multi-chip agent workloads.
Why it matters: Nvidia's real moat is software and its four million CUDA developers, not just silicon; a credible open Chinese alternative aimed at domestic chips is how that moat erodes, and it signals China's model makers and chipmakers closing ranks under export controls.
AMD buys Fei-Fei Li's World Labs for $8.2B to chase Nvidia on world models
AMD is acquiring World Labs for $8.2 billion, with founder Fei-Fei Li joining as executive vice president and chief scientist reporting to CEO Lisa Su. Founded in 2024, World Labs builds spatial-intelligence "world models"; its recent Atlas architecture predicts new camera views from 2D image inputs and, the company says, effectively solves the long-standing sparse-reconstruction problem in computer vision. The deal, expected to close by year-end pending regulatory approval, gives AMD an answer to Nvidia's open Cosmos world-model stack and a source of synthetic data for robotics simulation.
Why it matters: AMD is buying a frontier foundation-model team, not just talent, betting that owning spatial-intelligence models steers its chip roadmap and narrows Nvidia's ecosystem lead in robotics and simulation.
- AMD will acquire Fei-Fei Li's World Labs for $8.2 billion (TechCrunch AI)
- AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem (Latent Space (swyx))
- World Labs is Joining AMD (World Labs)
Google flies four TPUs to orbit October 1 in first Project Suncatcher test
Google will launch MVP, a refrigerator-size prototype satellite carrying four Trillium TPUs, on a SpaceX Falcon 9 from Vandenberg as part of the Transporter-18 rideshare. Built with Planet Labs, it draws about one kilowatt of solar power and aims to validate whether TPUs survive launch vibration, radiation, and vacuum cooling; Google says its Trillium chips already survived a proton-beam dose exceeding a five-year mission. The company estimates roughly 10,000 satellites would be needed to match a single 1-gigawatt terrestrial data center, and that launch costs must fall to around $200 per kilogram to make the economics work.
Why it matters: Orbital data centers are still a moonshot, but a real hardware test in space moves the idea from press release to measured failure points — and everyone from SpaceX to Blue Origin is chasing the same thing.
Qualcomm's Snapdragon 8 Elite Gen 6 runs a 30B MoE model on a phone
At its Snapdragon Summit, Qualcomm announced the Snapdragon 8 Elite Gen 6 and a higher-end Extreme Gen 6, both pitched around on-device AI. New sensing hubs can run models up to 200 million parameters continuously for a local voice-in/voice-out agent and speaker differentiation, while the Extreme variant can run a 30-billion-parameter mixture-of-experts model locally. Qualcomm noted the comparison to Apple's 20B MoE foundation model from WWDC. Motorola's Signature 27 will be the first device on the Extreme chip.
Why it matters: A 30B MoE running locally on a flagship phone pushes usable on-device inference well past the small-model tier, and gives app developers a real target for privacy-sensitive, offline agent features.
Alibaba unveils Qwen 4 and a new AI chip at Apsara, teases a multi-trillion-parameter model
At its Apsara conference Alibaba announced Qwen 4 and detailed a new in-house AI chip, while laying out plans to scale its flagship model. Reports indicate a planned model in the 5-trillion to 10-trillion-parameter range. Concrete specifications for both Qwen 4 and the accelerator remain thin, and much of the parameter detail comes from secondhand summaries rather than Alibaba's own materials.
Why it matters: Pairing a custom accelerator with multi-trillion-parameter ambitions is Alibaba's bid to blunt its Nvidia dependence and defend Qwen's lead among open Chinese models.
OpenAI says LLMs took its Jalapeño chip from concept to silicon in under 20 months
In an IEEE Spectrum account, OpenAI detailed how it used its own models to design Jalapeño, its 13.4-petaflop 4-bit inference accelerator with 232 GB of memory at 15.4 TB/s, claiming up to 3.6x lower end-to-end latency than Nvidia's GB300. A team of under 100, partnered with Broadcom, went from architecture to first silicon in under 20 months and from RTL to tape-out in nine, leaning on the open-source XLS high-level synthesis flow because it 'looks like software.' On one DeepSeek multi-head latent attention kernel benchmark, AI-written software climbed from 0.31% to 88.94% of the chip's theoretical ceiling in roughly 40 hours.
Why it matters: Concrete evidence that LLMs are compressing the front end of chip design — with the caveat that these are vendor-cited figures on OpenAI's own silicon, not independent benchmarks, and Broadcom did the physical implementation.
- How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip (IEEE Spectrum)
Huawei pulls Ascend 960DT forward to Q1 2027 — but the SuperPoD shrinks
At Huawei Connect, acting chairman David Wang said the next-generation Ascend 960DT AI chip is now expected in Q1 2027, moved up from a previously planned Q3, with claimed doubled performance. Huawei is pitching its Peerium Computing Architecture and UnifiedBus interconnect to lash chips into one machine; it says an Atlas 950 SuperCluster can link up to 256,000 accelerator cards. Analyst Rui Ma flagged a catch: the SuperPoD announced this week tops out at 4,096 chips, far below the 15,488-chip Atlas 960 SuperPoD Huawei had earlier described. So the chip arrives sooner, but the system around it is smaller than promised.
Why it matters: Huawei is the clearest test of whether export controls actually slow China's AI hardware. A pulled-in timeline days before the Trump–Xi meeting is as much signaling as engineering — and the shrunken cluster is the detail to watch.
Apple reportedly plans an M8 Ultra AI server, possibly with Nvidia NVLink
Apple is developing an enterprise server built on its own future M8 Ultra chips, in two- and four-chip configurations aimed at AI developers, businesses, and governments running inference on trained models, according to The Information. Apple is weighing Nvidia's NVLink Fusion to link the chips inside data centers. A launch wouldn't come before 2029 and the project could still be scrapped. It would be Apple's first server since it discontinued Xserve in 2011, and follows AI labs including OpenAI and Anthropic buying Mac minis and Mac Studios in bulk for AI workloads.
Why it matters: An Apple-silicon inference box borrowing Nvidia's interconnect would be a notable crack in the CUDA-and-x86 datacenter default, though the 2029 timeline and Apple's history of abandoning server hardware keep it firmly speculative.
Agility's Digit 5 is built to work fenceless next to people
Agility Robotics unveiled Digit 5, a humanoid designed to operate near workers without safety cages: it detects nearby people via AI and sensors and will stop, step aside, or squat to a seated position. It lifts up to 22.7 kg (40% more than Digit 4), charges in 9 minutes for 90 minutes of runtime, and is the first partner for Nvidia's Halos robotics safety platform. Agility says Digit 4 logged over 65,000 hours with GXO, Amazon, and Schaeffler; first Digit 5 deliveries start in early 2027.
Why it matters: Removing physical separation barriers is the practical unlock for putting humanoids on real warehouse and factory floors, and OSHA-style safety review is becoming a shipping requirement rather than a demo checkbox.
AWS benchmarks Blackwell G7 instances: native FP4 pays off on small MoEs
AWS published SageMaker benchmarks of its new G7 instances (NVIDIA RTX PRO 4500 Blackwell) against G5 (A10G) and G6 (L4) for 30B MoE inference. On a Qwen3-Coder-30B FP8 coding workload, ml.g7.12xlarge hit about 391 output tokens/second, 60.8% over G6 and 13% over G5, with lower P99 latency, using two GPUs and 64GB versus four GPUs and 96GB on the older families. G7 is the only generation with native FP4 tensor-core support, which AWS says gives it a structural edge on NVFP4-quantized MoE models; a separate Nemotron-3-Nano test put its cost per output token up to roughly 4.9x below G6.
Why it matters: For teams self-hosting small MoEs, Blackwell's native FP4 is starting to show up as concrete price-performance rather than spec-sheet headroom, though these are vendor numbers on the vendor's own hardware and regions.
- Benchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6 (AWS Machine Learning)
Alibaba Cloud, Cambricon and Ant Group join the PyTorch Foundation
At PyTorch Conference China in Shanghai, Alibaba Cloud and AI-chip designer Cambricon joined the Linux Foundation's PyTorch Foundation as Platinum members and Ant Group as Gold, alongside existing member Huawei. The pitch is device-agnostic PyTorch: Cambricon detailed hardening the framework's PrivateUse1 backend path for its MLU hardware and its work bringing Day-0 vLLM support to DeepSeek-V4 and GLM-5, while Huawei pushes native Ascend NPU support.
Why it matters: The people building on non-Nvidia silicon are now buying board seats to make PyTorch and vLLM portable across Chinese accelerators — a direct hedge against CUDA lock-in that could matter for anyone serving open models cheaply.
Preferred Networks lines up a 2028-2030 IPO to mass-produce MN-Core chips
Japan's Preferred Networks plans an IPO between 2028 and 2030 to fund volume production of its custom MN-Core inference processors, which the company claims run generative AI workloads up to 10x faster than conventional Nvidia GPUs — a figure it has not had independently verified. The chips use 3D-stacked memory to target bandwidth, the real inference bottleneck, and are fabbed by TSMC and Samsung. Backers include Toyota, SBI, Fanuc and NTT; the firm is valued around $2 billion with about 450 staff.
Why it matters: Another memory-bandwidth-first challenger to Nvidia on inference, though the headline 10x claim is the vendor's own and volume shipping is years out.
- Preferred Networks AI Chips IPO Targets Faster Generative AI Hardware (The Cryptonomist)
DeepSeek plans a 160,000-chip Huawei Ascend cluster for inference
DeepSeek intends to deploy at least 160,000 of Huawei's next-generation Ascend-950DT chips in an Inner Mongolia data center, according to Bloomberg as reported by The Decoder — which would be the largest known Huawei chip cluster. The chips would run inference only; DeepSeek still relies on Nvidia hardware for training. Huawei likely can't fulfill the full order for over a year given production and HBM memory shortages, though China's CXMT has begun small-batch HBM3E output while remaining several years behind Samsung, SK Hynix, and Micron.
Why it matters: It's a concrete measure of how far a leading Chinese lab can move inference off Nvidia — and how far it still can't, given the training gap and the memory bottleneck gating domestic accelerators.
Nvidia to buy Hugging Face for $12.9B
Nvidia agreed to acquire Hugging Face, the main distribution hub for open-weight models, for $12.93 billion — roughly $11.9 billion in purchase price plus up to $1 billion in staff retention stock. The deal is expected to close in the first half of 2027 pending regulatory approval. Jensen Huang pledged Hugging Face will stay an open, hardware-neutral platform where Nvidia compute is not required, and noted Nvidia is already its largest contributor with 500+ models and 250+ datasets. Hugging Face turned down a Nvidia investment at a $7 billion valuation just last year to stay independent.
Why it matters: The dominant chipmaker now owns the GitHub of open AI at a moment when big labs are designing their own silicon; every promise about neutrality and openness will be tested by regulators and the open-model community.
- NVIDIA to Acquire Hugging Face (NVIDIA)
- Nvidia buys Hugging Face, the GitHub of AI, for $13 billion (Ars Technica)
- Nvidia buys the front door to open AI as closed labs increasingly design their own silicon (The Decoder)
- Nvidia to spend $13 billion on Hugging Face, which will remain an open source platform (ABC News)
NVIDIA and CrowdStrike build a Nemotron-based agentic cyber defense
At Fal.Con 2026, CrowdStrike and NVIDIA unveiled SafeMind, an agentic cybersecurity system pairing CrowdStrike's models and harnesses with a defensive model built on open NVIDIA Nemotron and post-trained on CrowdStrike threat data. It runs an offensive red-team agent against a defensive blue-team agent in a continuous coevolution loop on a digital twin of NVIDIA's own network. CrowdStrike claims internal evals showed its Nemotron 3 Super-based 'Blue Solano' model beat leading frontier models on accuracy at 99% lower cost. A companion product, Falcon IQ, orchestrates more than 50 agents for assessment and remediation.
Why it matters: The pitch — post-train an open model on your own security data rather than rent a closed frontier API — is a concrete argument for why defenders may prefer inspectable open weights in high-stakes domains.
Apple accuses OpenAI of destroying evidence in trade-secrets suit
In a Monday filing supporting its motion for expedited discovery, Apple alleged that OpenAI is actively destroying evidence and that former iPhone engineer Chang Liu, now at OpenAI, both downloaded a confidential Apple circuit schematic and used it in his work. Apple says Liu retained access via a previously unknown authentication bug and enlisted an OpenAI colleague to help destroy evidence in June once he learned of the investigation. Apple is seeking a preliminary injunction to bar OpenAI from building hardware based on its technology and notes that more than 400 former Apple employees now work at OpenAI. OpenAI called the dispute a mess of Apple's own making and blamed residual-access mismanagement.
Why it matters: The case is now less about one engineer and more about how much of Apple's silicon know-how has walked into OpenAI's hardware effort, with an injunction on the table that could stall that program mid-flight.
AI labs are buying tens of thousands of Mac minis to train computer-use agents
The Decoder, citing The Information, reports OpenAI and rival labs have bought tens of thousands of Mac minis and Mac Studios to train computer-use agents on real desktop environments, with the most powerful configs sold out for months amid a memory-chip shortage. Anthropic is said to rent Mac minis through AWS. Apple's Mac revenue rose nearly 29% to $10.4 billion in the June quarter; software like Exo lets users cluster Macs to run large models locally.
Why it matters: Training agents to click through real GUIs means labs need real machines, not just GPUs — a reminder that the computer-use race runs on commodity desktop hardware, and that consumer supply is now colliding with frontier demand.
Anthropic's Model Hardware Standard lets agents drive lab gear
Anthropic introduced the Model Hardware Standard (MHS), a research-preview set of standardized drivers that let AI agents interface with physical devices through a common protocol within preset safety limits. In the first showcase, QuEra had Claude write and test a controller that restores a quantum computer's laser lock, recovering in 695 of 700 timed trials across seven fault types with no false success reports. Claude produced conventional software engineers could inspect and validate, rather than staying in the control loop at runtime.
Why it matters: MHS is Anthropic's bid to turn 'agents in the physical world' into a standard interface instead of a bespoke integration per rig — and the QuEra pilot is a rare concrete, independently verified deployment rather than a demo.
- Anthropic's new hardware standard lets AI agents control the physical world (Ars Technica AI)
- QuEra Uses Anthropic AI Agent to Automate Critical Quantum Computer Process (The Quantum Insider)
Hugging Face's $399 Microduck is an open-source robot you train with RL
Hugging Face and Pollen Robotics unveiled Microduck, a 25cm open-source bipedal robot priced at $399 and slated to ship before Christmas. It carries a camera, LiDAR, two IMUs and 15 actuators, and can waddle, grip up to 800g with its beak, self-right, crouch, and roller-skate; behaviors train in simulation and deploy directly to the hardware, with the SDK, simulator, and full RL training stack on GitHub. The launch comes as Hugging Face is reportedly set to be acquired by Nvidia.
Why it matters: A cheap, fully open sim-to-real loop is a more credible on-ramp to hobbyist embodied AI than another closed demo bot — and puts community-trained policies, not just canned behaviors, in reach.
- Hugging Face is selling a cute $399 open source duck robot, Microduck (TechCrunch AI)
- Microduck by Pollen Robotics & Hugging Face (r/LocalLLaMA)
OpenAI's Jalapeño inference chip beats Nvidia Blackwell in first benchmarks
At Hot Chips, OpenAI detailed Jalapeño, its first in-house accelerator, co-developed with Broadcom and built purely for LLM inference. On SemiAnalysis's InferenceX suite — verified in-lab but with numbers supplied by OpenAI — it claims 1.5x-1.9x more throughput per watt and 1.7x-3.6x lower latency than Nvidia's GB200/GB300 racks across GPT-OSS-120B, DeepSeek R1 and Kimi K2.5, all without speculative decoding. The chip taped out in November 2025, runs at 700W with HBM4, and is inference-only; it remains at engineering-sample stage with volume production not scheduled until 2027.
Why it matters: A first-generation ASIC out-performing Blackwell is unusual, and SemiAnalysis argues the fast software bring-up means 'the CUDA moat is potentially dead' — but the honest comparison is against Rubin, not Blackwell, where the two run roughly even on cost per token.
- OpenAI Jalapeño: Better Than Nvidia Blackwell (SemiAnalysis)
- OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's Blackwell and Rubin in inference benchmarks (The Decoder)
- OpenAI Jalapeno Custom AI ASIC at Hot Chips 2026 (ServeTheHome)
- OpenAI's Jalapeño chip is built for fast inference at scale, benchmarks show (TechCrunch AI)
- OpenAI Says Its New Chip Outperforms Nvidia's Blackwell As Nvidia Prepares Earnings Release (24/7 Wall St.)
Apple pitches the M5 Ultra Mac Studio as your local-inference escape hatch
Apple unveiled new Mac Studios with M5 Max and M5 Ultra (up to 512GB unified memory, 1.2TB/s bandwidth on the Ultra, a 50% jump over M3 Ultra), explicitly marketing on-device models 'without counting tokens.' A Forbes cost analysis finds the pitch shaky: a 128GB M5 Max breaks even against a $200/month subscription only in year three, and the open models that actually fit — Qwen3-Coder-Next 80B, gpt-oss-120b, Qwen3.8-27B — score well below Claude on agentic coding benchmarks, while frontier open weights like Kimi K3 (2.8T) don't fit at all. Prompt processing on Apple silicon also remains slow. Machines ship September 22.
Why it matters: The real case for buying the box is data sovereignty, not saving money — for anyone whose code or patient data legally can't leave the building, not a cheaper Claude.
Nvidia pushes Groq 3 LPX into production with a contested 4x-Cerebras claim
Nvidia moved the Groq 3 LPX — the inference accelerator from the Groq team it acquired for about $20B — into full production at Hot Chips. An Artificial Analysis benchmark clocked 3,400 tokens/sec on Gemma 4 31B at 100k context, which Nvidia frames as 4x faster than Cerebras's 882 tok/s. But the comparison hides the chip count: the SRAM-heavy LPU carries just 500MB each, so the result needs at least 64 accelerators to Cerebras's one or two, uses a dense best-case model, and ignores Cerebras's newer CS-4.
Why it matters: Token-generation speed is the currency of agentic workloads, but a headline that needs 64 chips to beat someone's one or two is a marketing number, not an efficiency one.
Nvidia lays out what it takes to run CUDA on RISC-V
At Hot Chips 2026, Nvidia detailed the requirements for extending CUDA to RISC-V host CPUs, per a Chips and Cheese writeup: an RVA23 server-class core, adherence to RISC-V's server SoC and platform specs, ACPI, PCIe coherency, and peer-to-peer PCIe. Nvidia is partnering with SiFive, which planned to demo a CUDA-capable RISC-V system at the conference. The bar is high enough that essentially no existing consumer RISC-V hardware qualifies, and wide ACPI support is likely years out.
Why it matters: A path for RISC-V CPUs to feed Nvidia GPUs would loosen the x86/Arm lock on AI host processors — but only for server-grade silicon, and not soon.
- Hot Chips 2026: CUDA Targets RISC-V (Chips and Cheese)
Cerebras CS-4 doubles throughput on the same WSE-3 die
Cerebras unveiled the CS-4, a rack-scale system still built on the 5nm WSE-3 chip but doubling CS-3 performance by pushing clock speed through more power and better cooling. A rack now holds three wafers instead of two and delivers up to 4,400 tokens/sec per user — claimed up to 30x faster than Nvidia GPU setups — with memory unchanged at 44GB per wafer. A modular 'Backpack' design and disaggregated inference via AMD and AWS Trainium round it out; SemiAnalysis views the networking gains as small. OpenAI already uses Cerebras for Codex Spark.
Why it matters: The gains come from brute clock scaling, not a new node — useful if you're latency-bound on agentic workloads and can actually get rack access.
Hosting Kimi K3 (2.8T) on 8 B300s: 92 tok/s at $190 per million tokens
A developer benchmarked Moonshot's 2.8T-parameter open-weight Kimi K3 on 8 B300 GPUs via Modal and vLLM (tensor parallel 8, native MXFP4): a 27-minute cold boot loading 1.56TB, ~0.9s TTFT, 92 tok/s steady decode, and about $190 per million output tokens — roughly $1,363/day kept warm. Unsloth's 1-bit UD-IQ1_S GGUF (594GB) ran on 8 A100-80GBs at ~9 tok/s but worked out 3.3x more expensive per token despite the cheaper hardware.
Why it matters: Concrete, reproducible economics for self-hosting a frontier-scale open model — and a reminder that extreme quantization can cost more per token than it saves once throughput collapses.
Memory shortage pushes Nvidia AI server prices up about 15%
Bloomberg reports that systems built on Nvidia's Vera Rubin and Grace Blackwell chips will cost 15%+ more for shipments early next year, driven by rising DRAM prices from Samsung, SK Hynix and Micron. Contract manufacturers have already warned customers including Microsoft, Google and Oracle. The bill lands on cloud giants and on labs like OpenAI and Anthropic that still depend on Nvidia even as they build their own silicon.
Why it matters: Training and inference capex just got more expensive at the hardware level — the kind of pressure that eventually flows downstream into API pricing and GPU availability.
Nvidia pays $6B for Poolside's model factory and 109 engineers
Per an investor letter first reported by Newcomer, Nvidia is licensing Poolside's "Model Factory" — the pipeline behind its Laguna model — extending job offers to 109 of Poolside's roughly 115 technical staff, and investing $1B at a $12B pre-money valuation, while the three founders stay on. Poolside frames it as "not an acquisition and not an acquihire" and plans to distribute the $6B to investors by the end of next year. Latent Space calls it a reverse-execuhire: unlike the Windsurf, Character and Scale deals where executives left and staff stayed, here the founders keep the shell to pivot while employees and investors cash out.
Why it matters: Nvidia builds its own Nemotron open models, so it is now buying model-building capability from a startup it also invested in — another deal structured to lock in tech and talent without a full acquisition, following Groq ($20B) and Enfabrica.
Four 2017 V100s match an RTX 5090 on Qwen3.8 decode via a hand-written FP4 translator
A developer got four Tesla V100s — Volta, with no native FP4 or FP8 silicon — to run Qwen3.8's published mixed NVFP4/FP8 weights at ~219 tok/s single-request decode, statistically tied with a 5090 running the NInfer engine at ~215 tok/s. The kernel, 'QPN', translates compressed weight fragments straight into Volta's FP16 tensor-core format while reading from HBM (hitting 71–82% of read bandwidth) and maps a k=7 speculative-decode round onto Volta's native 8-row tile. Caveats are real: four GPUs versus one, ~4x slower prefill, and ~A$600 for the cards alone (loud, power-hungry datacenter hardware).
Why it matters: The takeaway is that a lot of 'too old for AI' datacenter hardware is missing software, not silicon — useful ammunition for anyone pricing out cheap self-hosted inference.
Unitree's $50B IPO runs on a circular robot-data economy
Unitree Robotics hit around $50B in its Shanghai debut, closing up 460% and becoming the first humanoid maker to list on the mainland. Per the FT, much of the demand is circular: state-backed training centers buy the robots, teach them tasks via teleoperation, then sell the collected data back to the manufacturers — nearly three-quarters of Unitree's humanoid revenue came from education and research. Analysts question both the 35x-revenue valuation and the data's usefulness, with one center manager saying only two to three of every eight training hours are usable.
Why it matters: China now has its own version of the circular-financing critique aimed at US AI firms — a reason to discount headline humanoid-robot demand before extrapolating it.
- China now has its own AI circular financing scheme (The Decoder)
RAMageddon: memory prices up 500% in a year, 128GB DDR5 now $3,399
Per Tom's Hardware via Latent Space, DRAM prices have climbed roughly 500% in 12 months and up to 10x their lowest-ever tracked levels, with 128GB DDR5 kits hitting $3,399. Hyperscalers have reportedly locked in almost all of 2027's global DRAM capacity with advance deposits, and mainstream DRAM now sells for over half the price of gold by weight. Moore's Law, at least for memory, has gone into reverse.
Why it matters: For anyone building local-inference rigs or spec'ing self-hosted deployments, the cost math just broke — high-RAM boxes for large MoE models are suddenly a luxury, not a weekend upgrade.
- [AINews] Memory prices up 500% in 12 months (Latent Space (swyx))
- Memory prices climb 500% in 12 months, up to 10x the lowest ever tracked prices - 128GB of DDR5 now $3,399 (r/LocalLLaMA)
Nvidia backstops $105B for OpenAI's record Ohio data center
OpenAI signed a 20-year lease for the PORTS-Pike campus in Pike County, Ohio, built and owned by SoftBank's SB Energy on a decommissioned uranium-enrichment site. Nvidia becomes the exclusive chip supplier and guarantees up to $105B of the finished facilities' residual value on the first 4.25 IT-GW (of 8 IT-GW total, backed by ~10 GW of new gas generation), plus a $1.5B stake in SB Energy; first 800 MW is slated for 2028. The Wall Street Journal notes nine tech firms now carry roughly $3 trillion in mostly-AI commitments off their balance sheets.
Why it matters: Nvidia is now simultaneously OpenAI's supplier, investor and loan guarantor — Jensen Huang insists 'OpenAI will pay the lease,' but the structure is the clearest test yet of whether AI's circular financing holds up if demand doesn't fill the racks.
- NVIDIA Guarantees SB Energy's PORTS-Pike Technology Campus in Ohio to Exclusively Host NVIDIA AI Compute (NVIDIA Newsroom)
- OpenAI signs record Ohio data center lease with Nvidia backing up to $105 billion (The Decoder)
- OpenAI announces massive data center in Ohio with Nvidia guarantee (Axios)
- Nvidia investing $1.5B in SoftBank data center developer behind OpenAI project (TechCrunch AI)
- Nvidia backs $105B for OpenAI's mega data center (The Neuron)
Groq raises $350M at $3.5B, half its old valuation, and leans into Nvidia clouds
Groq raised $350M led by Disruptive, with planned Nvidia participation, at a $3.5B valuation — down from $6.9B last September, after Nvidia hired founder Jonathan Ross and top talent in a $20B licensing deal. The company insists it isn't a down round but a reset for the 'post-Nvidia-licensing-deal' Groq, which has pivoted from building its own LPU inference chips to operating Nvidia systems as a neocloud. It now runs 13 data centers and plans to scale from 54 MW to 200+ MW in 2027.
Why it matters: A one-time custom-silicon challenger now reselling Nvidia GPUs is a blunt signal about how hard it is to compete on inference chips — and neocloud economics (capex, debt, fast-depreciating hardware) remain unproven.
- Groq raises $350M to fuel its pivot from AI chips to neocloud (TechCrunch AI)
When the AI-companion startup folds, the kid's robot dies
MIT Technology Review traces Moxie, the $800 AI robot marketed as a social-skills companion for neurodivergent children, through two corporate collapses that bricked the cloud-dependent device. When maker Embodied shut down in 2024, an engineer shipped OpenMoxie, open-source firmware to keep the robots running locally, but many families could not migrate before the servers went dark; a second owner then folded in 2025. The piece is a case study in the planned obsolescence of emotionally-bonded, always-online consumer AI hardware, and the thin clinical evidence behind therapeutic robots.
Why it matters: Any product that offloads its brain to a startup's servers inherits that startup's runway. 'The company folded' is now a failure mode for a child's best friend.
- What happens when a kid's robot best friend dies? (MIT Technology Review)
US to allies: join our AI bloc or China's, not both
A draft State Department letter reviewed by Reuters would tell the 35 signatories of Washington's June 'AI Opportunity Statement' that membership in its Pax Silica initiative — covering AI models, semiconductors, and critical minerals — 'cannot be held alongside' China's rival World Artificial Intelligence Cooperation Organization. Kazakhstan, a critical-minerals supplier that joined both frameworks, is the early test case. The stated aim is to choke China's access to the inputs needed for frontier AI.
Why it matters: Export-control lines are hardening into full ecosystem exclusivity: where a model's weights, chips, and minerals come from is becoming a diplomatic loyalty test that will shape who can build and deploy what.
- The U.S. is drawing a line in the global AI race with China (calcalistech.com)
- AI's New Red Flag: US Dangles Pax Silica At Partners To Sideline China (International Business Times)
OpenAI puts GPT-5.6 Sol on Cerebras for 750 tokens per second
OpenAI opened a limited preview of Ultrafast mode, a Responses API tier that runs GPT-5.6 Sol at up to 750 output tokens/sec — roughly 14x standard — powered by Cerebras' wafer-scale engines rather than GPUs. Cerebras claims it cleared all 2,500 Humanity's Last Exam questions in 11 hours versus 78 for Claude Fable 5, at comparable accuracy, and a 5.6x end-to-end speedup on GDP-Val. Access is gated to a small customer set for now, aimed at incident response, trading and security workloads.
Why it matters: If frontier-quality output at 750 tok/s holds up outside vendor benchmarks, latency-bound agent loops stop being a reason to drop down to a smaller model.
Nvidia guarantees its own chips' resale value to unlock $500B in AI debt
Nvidia signed letters of intent with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to mobilize over $500 billion in third-party capital for data centers, fabs, and power plants. To make the financing pencil out, Nvidia will backstop up to 25% of the residual value of its own installed GPUs on a per-project basis, effectively absorbing part of the depreciation risk. Jensen Huang argues the hardware lasts far longer than critics claim, citing A100s still earning revenue six years on and H100 rental rates rising from $1.70 to $2.35 per GPU-hour. The move reads as a direct rebuttal to Michael Burry's warning that GPU depreciation is understated by ~$176B through 2028.
Why it matters: The whole AI buildout rests on how long a GPU stays economically useful. Nvidia putting its balance sheet behind that number, rather than just selling chips, is a tell about how circular the financing has become, and how much rides on utilization staying high.
AMD buys Taalas to etch whole models into silicon
AMD acquired chip startup Taalas, which builds model-specific integrated circuits that hard-wire a model's weights into silicon rather than loading them onto general-purpose GPUs. Early demos claim up to 17,000 tokens per second on these etched-model chips. AMD is framing it as an enterprise inference play, betting the market goes vertical as serving costs dominate.
Why it matters: If per-model ASICs deliver order-of-magnitude throughput, the economics of inference shift away from flexible GPU fleets toward fixed silicon per model, changing how anyone plans a serving stack for the next few years.
- AMD acquires Taalas to boost inference performance by etching models in silicon (The Register)
- [AINews] AMD buys Taalas (Latent Space (swyx))
- AMD Acquires Taalas to Advance Compute Solutions for AI Inference (r/LocalLLaMA)
SK Hynix's $476K bonus is bleeding Samsung's chip engineers dry
SK Hynix's record HBM profits translated into a roughly $476,000 per-employee cash bonus this year, versus about $135,000 for Samsung's loss-making foundry division, and Samsung engineers are defecting en masse. A union survey found 81.5% of foundry staff want out within two years; Samsung won an 18-month injunction blocking two former workers from joining its rival. The exodus threatens Samsung's one structural edge in HBM4: being the only memory maker that also runs its own advanced logic foundry.
Why it matters: The AI boom's constraint is shifting from GPUs to the HBM stacked on them, and whoever retains the memory-and-logic talent controls the supply that feeds every Nvidia accelerator.
- Samsung's chip workers are jumping ship to rival SK Hynix (MIT Technology Review)
Chinese DRAM maker CXMT surpasses Intel's market cap on a 500% debut
CXMT, mainland China's only integrated device manufacturer mass-producing general-purpose DRAM, surged nearly 500% on its first trading day to roughly RMB 3.28 trillion, the largest company by value on China's A-share market. That edges past Intel, which closed the prior day at about $465.6 billion (~RMB 3.15 trillion). The Hefei-based firm is central to China's push for domestic memory supply.
Why it matters: Memory is the bottleneck for AI accelerators; a well-capitalized domestic DRAM champion signals China intends to close the HBM and DRAM gap that export controls were meant to hold open.
Anthropic asks SK Hynix for supplies to build its own chips
SK Group chair Chey Tae-won said Anthropic approached SK Hynix, one of the largest memory makers, for supplies to make its own semiconductors, speaking on stage alongside Dario Amodei at a San Francisco AI event. Chey called it remarkable for an AI developer to pursue its own silicon. The visit coincided with South Korea's president convening an AI summit, where Nvidia also announced partnerships with Naver and SK Group.
Why it matters: After committing to 2GW of AMD MI450s last week, Anthropic sniffing at custom silicon signals it wants leverage over both the Nvidia and AMD supply queues.
Etched raises $300M at $10.3B to build transformer-inference systems
Etched closed a $300M Series C at a $10.3B valuation led by Sequoia, with a16z, SK Hynix, and Jane Street participating, doubling its December valuation in seven months. The company says it has already booked $1B in orders and is shipping full rack systems, not just chips, with a low-voltage prefill chip and a 'cluster-scale memory' interconnect for the decode phase. It pushes back on the perception that its silicon runs only specific LLMs, claiming support for MoE models and non-transformer designs like Mamba. Etched also opened an 80,000 sq ft, 10 MW facility in Milpitas, framing its pitch as 'run the world's inference.'
Why it matters: Inference-specialized silicon is graduating from thesis to booked revenue, and the more credible these alternatives get, the more pricing pressure Nvidia faces on the serving side.
Anthropic commits to 2GW of AMD MI450 GPUs; AMD invests up to $5B
AMD will invest up to $5 billion in Anthropic, which in turn will deploy up to 2 gigawatts of Instinct MI450-series accelerators in Helios rack systems — MI455X GPUs paired with EPYC "Venice" CPUs, Pensando networking and ROCm — with the first gigawatt landing in H1 2027. AMD's stake is milestone-gated on deployment, echoing its 6GW OpenAI and 6GW Meta arrangements. A multi-year engineering program will use Claude to improve AMD's ROCm software, and AMD will run Claude internally across its dev teams.
Why it matters: It's another circular chip-lab financing loop, but it gives Anthropic a real second GPU source alongside Nvidia, Amazon Trainium and Google TPUs — and puts Claude to work hardening the weakest part of AMD's stack, its software.
Google reportedly bakes Gemini's architecture into 'Frozen v2' silicon
Per The Information, Google is building a server chip internally called Frozen v2 that hardcodes parts of Gemini's model architecture (not its weights) directly into hardware, claiming 6-to-10x more tokens per watt than its current TPUs, with deployment targeted for 2028. New weights can still be loaded, so the chip survives model updates; an earlier Jeff Dean design that froze weights themselves was scrapped as too brittle. It is meant for internal inference only, and the report nudged Alphabet stock up about 3% ahead of earnings.
Why it matters: Inference margin is the new competitive front, and specializing silicon to a single architecture is the logical extreme of the efficiency race, at the cost of being locked to that architecture.
First loan backed by inference chips: $400M for SambaNova silicon
AI inference cloud General Compute landed a $400M loan from Upper90, reportedly the first financing to use inference-specific chips as collateral — SambaNova's power-efficient SN50, which the startup claims runs 16x faster than GPU clouds. Upper90 pioneered GPU-backed lending with Crusoe in 2021; it's now betting the next wave is cheap inference for open models, outside Nvidia's ecosystem.
Why it matters: Capital markets are beginning to price non-Nvidia inference silicon as a financeable asset, a small crack in Nvidia's dominance and a signal that serving open models cheaply is becoming its own infrastructure category.
OpenAI's actual first device is a $230 light-up keyboard for Codex
Days after reports of a screenless smart speaker, OpenAI's first branded hardware turned out to be the Codex Micro — a $230, 13-key mechanical keypad built with Work Louder and sold through OpenAI's merch store. Its RGB 'Agent Keys' show live status for up to six Codex threads (thinking, done, needs input, error), with a rotary dial to set an agent's reasoning level and a joystick to launch workflows. It's a limited run, ships via Bluetooth/USB-C around July 24, and is explicitly positioned as a novelty 'command center' for managing fleets of coding agents.
Why it matters: It's a gimmick, not the Jony Ive companion device — but the hardware design encodes a real workflow assumption: developers now juggle enough parallel agents that they need an ambient dashboard to see which one is stuck.
OpenAI's first device: a screenless speaker built to feel alive
Bloomberg reports OpenAI's debut hardware product is a portable, screenless smart speaker pitched internally as a 'new type of home computer for the AI era.' It pairs a camera and sensors with the just-launched GPT-Live voice mode, and adds mechanical parts that physically move to make it seem lifelike. Unveiling is planned for later this year with a 2027 release; Apple's trade-secrets suit over hardware chief Tang Tan could delay it. It is reportedly the first of about five devices, including a phone replacement, a pendant, and home robotics.
Why it matters: A camera-equipped, always-listening, deliberately anthropomorphized device with access to your email is a very different threat model than a chatbot tab — and the same GPT-4o sycophancy that caused problems now ships with a motor.
- OpenAI's first hardware product is a screenless AI speaker designed to feel alive (The Decoder)
- OpenAI's First Device Will Be Movable, Screenless Speaker Built as AI Companion (Bloomberg.com)
- OpenAI's first hardware device is reportedly a screenless speaker that can move (TechCrunch AI)
- OpenAI's first hardware device will be a speaker, Bloomberg News reports (Reuters)
Apple's OpenAI complaint: 400 poached staff, an auth bug, and prototypes at interviews
Details from the 41-page filing sharpen the case first reported last week: Apple says 400+ ex-employees now work at OpenAI, that engineer Chang Liu exploited a 'rare' authentication bug to reach Apple's network weeks after leaving ('LOL, I found out I can access the [network storage]'), and that hardware chief Tang Tan had candidates bring CAD files and physical prototypes to interviews. Apple also alleges io used its confidential metal-finishing techniques by misleading a supplier. OpenAI: 'We have no interest in other companies' trade secrets.'
Why it matters: Strip the espionage framing and this is a talent-mobility fight — OpenAI is well-funded enough to ignore the Valley's no-poach norms, and discovery could set precedent for how AI labs recruit from incumbents.
- The wildest allegations in Apple's trade secrets lawsuit against OpenAI (TechCrunch AI)
- These are the wildest claims in Apple's lawsuit against OpenAI (Fortune)
- Apple sues OpenAI after ex-engineer allegedly used bug to steal trade secrets (Ars Technica AI)
- OpenAI is breaking Silicon Valley's unwritten code. That's why Apple is so angry. (Business Insider)
SK Hynix's StreamDQ moves weight dequantization into HBM
An SK Hynix paper proposes StreamDQ, a near-memory architecture that performs on-the-fly weight dequantization inside custom HBM for high-throughput, large-batch LLM inference. It reports up to 7.08x speedup and 90.23% lower energy on mixed-precision GEMM.
Why it matters: If dequantization happens in the memory subsystem rather than the GPU, quantized serving stops paying the bandwidth tax on every weight fetch — potentially a big lever for FP4 and mixed-precision inference at scale.
- Near-memory Dequantization Architecture In Custom HBM for LLM inference (SK hynix) (Semiconductor Engineering)
Caltech spinout claims a full 27B model running on an iPhone
PrismML, a Khosla-backed Caltech spinoff, says it compressed Alibaba's Qwen 3.6 27B from ~54GB to under 4GB and got it running on an iPhone 17 Pro, with open weights due next Tuesday. Crucially, it claims all 27B parameters stay active, versus Apple's own new on-device model that uses a sparse 20B architecture with only 1-4B active at a time. CEO Babak Hassibi says the technique shrinks models 'without hindering performance,' the usual claim that a benchmark will need to settle.
Why it matters: If the quality claim survives contact with real evals, a genuinely dense 27B on a phone changes the on-device ceiling from toy assistants to something that can run agents and code. Weights next week means the community can check the math fast.
$80 Tesla P100s ran silently noisy math in llama.cpp for years; a 3-line patch fixes it
A years-old llama.cpp CUDA bug forced the Pascal P100 (sm_60) down an fp16 math path that the GTX 10-series and P40 (sm_61) were long ago exempted from. Measured against fp32-reference logits on Qwen3.6-27B, the fix cut median KL divergence ~2300x (0.0023 to 0.000001) and lifted top-token agreement from 96.5% to 99.9% — with decode ~1.4% faster, since real workloads are GEMM/bandwidth-bound, not fp16-vector-bound. The patch simply extends the sm_61 exemption to sm_60; it's shipped in a turboquant fork because GGML bans AI-assisted contributions, and the bug was isolated by an agent loop running Fable 5.
Why it matters: P100s are ~$80 with 16GB HBM2 at 732 GB/s amid a DRAM crunch; a chunk of their reputation for 'worse' output was this bug, and the fix is measured only on sm_60 — not the all-GPUs panic some will read into it.
Apple sues OpenAI, alleging a 'coordinated campaign' to steal hardware secrets
Apple filed suit in California federal court accusing OpenAI of a systematic effort to misappropriate trade secrets for its unreleased devices, naming hardware chief Tang Tan (ex-iPhone/Watch design lead) and former engineer Chang Liu. The complaint says 400+ ex-Apple staff now work at OpenAI, that Liu downloaded dozens of confidential hardware files on an Apple laptop he never returned, and that Tan told candidates to bring 'actual parts' to interviews. OpenAI denies any interest in others' trade secrets; io Products, the Jony Ive startup OpenAI bought for ~$6.5B, is also a defendant.
Why it matters: The 2024 ChatGPT-in-iOS partnership has fully collapsed into a talent-and-IP war, and the timing — with OpenAI's device slipping to 2027 and an IPO rumored — makes this more than a spat over departing engineers.
- Apple sues OpenAI for allegedly running a "coordinated campaign" to steal trade secrets through poached employees (The Decoder)
- Apple files lawsuit accusing ChatGPT maker OpenAI of stealing trade secrets (AP News)
- Apple accuses OpenAI of using stolen trade secrets to create its upcoming AI gadgets in new lawsuit (CNN)
- Apple Sues OpenAI, Accusing It of Stealing Company Secrets (The New York Times)
- Apple sues OpenAI for trade secret theft (Axios)
SK Hynix raises $26.5B in the largest-ever foreign US IPO
The HBM memory maker sold 177.9M ADRs at $149 each on Nasdaq, raising $26.5B — topping Alibaba's 2014 record — with demand reportedly 7x oversubscribed and the stock opening 14% above price. Proceeds fund a new Korean fab, a packaging plant and EUV scanners to ease the AI-driven memory shortage. Commerce Secretary Lutnick is separately pressing SK Hynix and Samsung to build US fabs, while Micron pledged $250B in domestic manufacturing.
Why it matters: HBM is the real bottleneck behind every GPU order; a supplier flush with $26.5B and under US pressure to onshore is a signal about where inference capacity — and its cost — goes next.
ZML's LLMD promises peak inference across Nvidia, AMD, TPU, Apple and Intel
Paris startup ZML, backed by Yann LeCun, launched LLMD, an inference server that runs open-source LLMs at (claimed) maximum speed across Nvidia, AMD, Google TPU, Apple Metal and Intel Arc silicon. The pitch is breaking vendor lock-in and letting shops mix cheaper or lower-power chips; ZML says it's co-designing silicon with European chipmakers like Axelera, SiPearl and VSORA. LLMD is free but not open source, launched to gather usage data. The 20-person team has raised ~$20M and enters a crowded field against vLLM, SGLang and Baseten.
Why it matters: A genuinely chip-agnostic inference layer would loosen Nvidia's grip and give infra teams real leverage on cost-per-token — if the cross-vendor performance claims survive independent benchmarks.
Qualcomm launches GenieX to run LLMs on Snapdragon Windows laptops
Qualcomm, late to the on-device SDK race, released GenieX for running LLMs across CPU, GPU, and NPU on its Windows laptops. Early hands-on reports: ~20 tok/s on Gemma 4 26B (A4B) with 0.5s to first token on GPU/NPU, and ~10 tok/s for Qwen 3.6 27B with MTP on GPU. Standard Q4_0 GGUFs reportedly run via llama.cpp on the CPU.
Why it matters: Usable NPU/GPU offload on mainstream Windows laptops widens the hardware base for local inference beyond Apple Silicon and discrete NVIDIA cards — if the tooling holds up in practice.
Anthropic in early talks with Samsung to build a custom AI chip
The Information reports Anthropic is exploring a custom processor built on Samsung's 2nm process and advanced packaging, and has hired Clive Chan, an early member of OpenAI's silicon team. The project is very early: no design, testing, or defined function yet, and Anthropic insists Nvidia GPUs, Google TPUs, and AWS Trainium will remain central. Samsung, SK Hynix, and Micron were strategic investors in Anthropic's $65B Series H. The move follows OpenAI's Broadcom-built 'Jalapeño' inference chip unveiled last week.
Why it matters: Every frontier lab now wants leverage over Nvidia and its own performance-per-watt story; the question is whether Anthropic can ship silicon years behind Google and Amazon without derailing its rented-compute supply lines.
Anthropic in early talks with Samsung for a custom AI chip
The Information reports Anthropic is discussing a custom accelerator with Samsung, though workloads, performance targets and process node are all undecided. Samsung offers its 4nm node and a data-center-tuned 2nm SF2P process entering production this year. Anthropic told press that AWS, Google and Nvidia silicon remains central to its strategy, and it has hired chip engineers including Clive Chan, an early member of Tesla's and OpenAI's silicon teams.
Why it matters: It follows OpenAI's Broadcom-built 'Jalapeño' inference chip by days: every major lab now wants custom silicon to escape Nvidia margins and control inference cost-per-watt. Whoever runs inference cheapest keeps more revenue.
- Anthropic in Talks With Samsung to Manufacture Custom AI Chip (The Information)
- Anthropic reportedly in talks with Samsung to manufacture custom AI chip (SiliconANGLE)
- Anthropic is discussing a new custom chip with Samsung (TechCrunch)
- Anthropic reportedly explores custom chip manufacturing with Samsung while insisting Nvidia still matters (The Decoder)
- Samsung seen in talks to manufacture custom AI chips for Anthropic (The Korea Economic Daily)
Meituan's LongCat-2.0: 1.6T params trained entirely on domestic chips
Meituan open-sourced LongCat-2.0, a 1.6-trillion-parameter model with a 1M-token context window, and claims it is the first trillion-parameter model to complete both pre-training and inference on a ~50,000-card domestic cluster of AI ASIC superpods. That goes a step beyond DeepSeek-V4-Pro, which Meituan says used home-grown chips only for inference. Pre-training is the far more compute-intensive phase, making the claim notable if it holds up.
Why it matters: If verified, it signals Chinese accelerators can handle frontier-scale training, not just inference — eroding one of the assumptions behind US export controls.
- Meituan claims China's biggest AI model trained on local chips (South China Morning Post)
Chip geopolitics: Korea's $1T bet, Taiwan raids Super Micro
South Korea committed $1 trillion across memory-chip production, AI data centers and humanoid robots, with President Lee calling semiconductors, physical AI and data centers the triple axis for a great leap forward. The same day, Taiwanese prosecutors raided Super Micro offices and partner firms over alleged smuggling of Nvidia AI chips into China; Super Micro's stock fell 8% and a co-founder was reportedly indicted.
Why it matters: The hardware supply chain is now an explicit instrument of state policy — both massive subsidies and criminal enforcement — and that volatility flows straight through to GPU and memory prices developers pay.
Samsung and SK Hynix commit ~$518B to new chip hub for AI demand
Samsung and SK Hynix, backed by the South Korean government, will invest a combined 800 trillion won (~$518B) in a new chipmaking hub in the country's southwest, with each building two fabs; The Decoder puts the total program nearer $590B including packaging and next-gen chip spending. The two firms control roughly 80% of the high-bandwidth memory market AI workloads depend on. Jefferies expects memory prices to rise 40-50% in Q3 2026 and another 30-40% in Q4, with relief unlikely before 2028.
Why it matters: HBM and DRAM price spikes are already pushing up hardware costs (Apple has hiked Mac prices), so anyone budgeting GPU or local-inference builds should expect memory to stay expensive into 2027.
Everyone wants off Nvidia: OpenAI's Jalapeño joins the custom-silicon rush
OpenAI detailed Jalapeño, a custom inference chip built with Broadcom, joining Google, Apple, and SpaceX in building their way out of single-supplier risk. The framing is hedge, not clean break — more control and hardware tuned to specific workloads, echoing Apple's gains from dropping Intel. The same discussion noted Groq raising $650M after Nvidia poached its top talent.
Why it matters: Custom inference silicon from the largest API providers could reshape pricing and availability downstream. If Jalapeño lands, it's another lever OpenAI gains over the cost curve that determines what you pay per token.
PyTorch's TokenSpeed-kernel makes multi-silicon inference a registry problem
A PyTorch blog details TokenSpeed-kernel, a standalone kernel subsystem that decouples the inference runtime from hardware-specific code via a public API (mha_prefill, moe_apply, etc.) plus a registry-and-selector that dispatches to platform kernels. Using GPT-OSS 120B on AMD MI355X (CDNA4) as the test case, Gluon-backed attention and MoE kernels delivered 1.6–3.6x end-to-end throughput over the portable Triton path, with the AMD kernels published separately as tokenspeed-kernel-amd and already adopted by vLLM. NVIDIA Blackwell paths sit behind the same API via FlashInfer/TensorRT-LLM wrappers.
Why it matters: Backend selection leaking into model code is a real maintenance tax as GPU vendors, quant formats, and architectures multiply. A clean kernel boundary that vLLM can borrow is how AMD stays a first-class inference target rather than a perpetual afterthought.
OpenAI and Broadcom tape out 'Jalapeño,' a custom LLM inference chip
OpenAI unveiled Jalapeño, its first custom accelerator (an 'Intelligence Processor') built with Broadcom specifically for LLM inference, with OpenAI doing chip design and Broadcom contributing silicon and Tomahawk networking. OpenAI claims design-to-tape-out took nine months — partly accelerated by its own models — and 'substantially better' performance per watt, though these are self-reported numbers with no technical report yet. Engineering samples are already running GPT-5.3-Codex-Spark in the lab; large-scale deployment is planned for late 2026 at gigawatt scale, with Microsoft reportedly committed to buying 40% of the first run. Community reverse-engineering pegs it as TPU-like, roughly 216GB HBM3E and ~10 PFLOPS FP4.
Why it matters: If the perf-per-watt claims hold, OpenAI gains leverage over inference economics and its Nvidia dependence — but until an independent technical report lands, treat the numbers as marketing.
- OpenAI and Broadcom announce chip designed for LLM inference at scale (Ars Technica AI)
- OpenAI and Broadcom unveil "Jalapeño," a custom chip built for LLM inference (The Decoder)
- OpenAI unveils its first custom chip, built by Broadcom (TechCrunch AI)
Qualcomm enters the data center with Dragonfly C1000 and buys Modular for ~$4B
Qualcomm announced the Dragonfly C1000, a data-center processor optimized for AI agents and low power, with Meta planning to deploy it starting 2028. Alongside it, Qualcomm is acquiring Chris Lattner's Modular — maker of the cross-architecture Mojo/inference stack — for roughly $4 billion, with Modular saying Mojo open-sourcing stays on track. Qualcomm nearly doubled its non-smartphone revenue forecast to $40B by 2029 (targeting $15B from data centers); the stock jumped 15% after hours.
Why it matters: The Modular buy gives Qualcomm a serious CUDA-alternative software story to pair with its silicon — another front in the slow erosion of Nvidia's lock-in.
Seven Chinese vendors are now shipping H100/H200-class accelerators
A widely-shared LocalLLaMA writeup maps at least seven Chinese AI-chip makers shipping today: 'three dragons' (Huawei Ascend, Alibaba T-Head, Baidu Kunlunxin) and 'four snakes' that mostly IPO'd in the last six months (MetaX, Moore Threads, Biren, Iluvatar CoreX). Current parts land around H100, next-gen targets H200, and production is shifting from TSMC to SMIC. The post cites a CHITEX talk for many specifics and flags vendor/analyst figures as unverified. NVIDIA's China GPU share reportedly fell from 95% to 55% in two years. Separately, a Chinese supercomputer reclaimed the world's-fastest spot for the first time since 2017.
Why it matters: Chinese open-weight models (Qwen, DeepSeek, GLM) are increasingly co-designed with domestic silicon, with its own form factor, interconnect, and HBM. If you run open weights, the hardware you target in two years may not be NVIDIA.
Reflection rents $6.3B of GB300s from SpaceX, the third neocloud deal
Open-weight lab Reflection AI will pay SpaceX $150M/month from July 2026 through 2029 for immediate access to Nvidia GB300 chips at the Colossus 2 data center near Memphis — a deal worth up to $6.3B, with a 90-day exit clause. It is smaller than SpaceX's Anthropic ($1.25B/month) and Google ($920M/month) contracts. Tallied together, SpaceX's GPU rentals annualize to roughly $28B/year at implied Blackwell pricing above $10/hour, about twice CoreWeave's current revenue.
Why it matters: SpaceX has quietly become a major 'neocloud,' and GPU brokerage is emerging as a strategic layer between model builders and hardware supply — with Reflection pitching open weights as the hedge against closed-model access being revoked.
- SpaceX inks compute deal with Reflection AI, an open source AI lab (TechCrunch AI)
- [AINews] SpaceX is already a $28B/yr Neocloud (Latent Space (swyx))