<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><title>gonioAI — Hardware &amp; chips</title><link>https://gonioai.pages.dev/topics/hardware/</link><description>Hardware &amp; chips stories from gonioAI.</description><language>en</language><lastBuildDate>Tue, 11 Aug 2026 10:45:13 +0000</lastBuildDate><item><title>Nvidia guarantees its own chips' resale value to unlock $500B in AI debt</title><link>https://the-decoder.com/nvidia-guarantees-its-own-chips-value-to-unlock-500-billion-in-ai-infrastructure-financing</link><guid isPermaLink="false">2026-08-11:hardware:https://the-decoder.com/nvidia-guarantees-its-own-chips-value-to-unlock-500-billion-in-ai-infrastructure-financing</guid><pubDate>Tue, 11 Aug 2026 07:00:00 +0000</pubDate><description>Nvidia signed letters of intent with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to mobilize over $500 billion in third-party capital for data centers, fabs, and power plants. To make the financing pencil out, Nvidia will backstop up to 25% of the residual value of its own installed GPUs on a per-project basis, effectively absorbing part of the depreciation risk. Jensen Huang argues the hardware lasts far longer than critics claim, citing A100s still earning revenue six years on and H100 rental rates rising from $1.70 to $2.35 per GPU-hour. The move reads as a direct rebuttal to Michael Burry's warning that GPU depreciation is understated by ~$176B through 2028.

Why it matters: The whole AI buildout rests on how long a GPU stays economically useful. Nvidia putting its balance sheet behind that number, rather than just selling chips, is a tell about how circular the financing has become, and how much rides on utilization staying high.</description></item><item><title>AMD buys Taalas to etch whole models into silicon</title><link>https://www.theregister.com/systems/2026/08/06/amd-acquires-ai-chip-startup-taalas-to-boost-inference-performance-by-etching-models-into-silicon/5284344</link><guid isPermaLink="false">2026-08-07:hardware:https://www.theregister.com/systems/2026/08/06/amd-acquires-ai-chip-startup-taalas-to-boost-inference-performance-by-etching-models-into-silicon/5284344</guid><pubDate>Fri, 07 Aug 2026 07:00:00 +0000</pubDate><description>AMD acquired chip startup Taalas, which builds model-specific integrated circuits that hard-wire a model's weights into silicon rather than loading them onto general-purpose GPUs. Early demos claim up to 17,000 tokens per second on these etched-model chips. AMD is framing it as an enterprise inference play, betting the market goes vertical as serving costs dominate.

Why it matters: If per-model ASICs deliver order-of-magnitude throughput, the economics of inference shift away from flexible GPU fleets toward fixed silicon per model, changing how anyone plans a serving stack for the next few years.</description></item><item><title>SK Hynix's $476K bonus is bleeding Samsung's chip engineers dry</title><link>https://www.technologyreview.com/2026/07/28/1140853/samsung-chip-workers-exodus-sk-hynix</link><guid isPermaLink="false">2026-07-28:hardware:https://www.technologyreview.com/2026/07/28/1140853/samsung-chip-workers-exodus-sk-hynix</guid><pubDate>Tue, 28 Jul 2026 07:00:00 +0000</pubDate><description>SK Hynix's record HBM profits translated into a roughly $476,000 per-employee cash bonus this year, versus about $135,000 for Samsung's loss-making foundry division, and Samsung engineers are defecting en masse. A union survey found 81.5% of foundry staff want out within two years; Samsung won an 18-month injunction blocking two former workers from joining its rival. The exodus threatens Samsung's one structural edge in HBM4: being the only memory maker that also runs its own advanced logic foundry.

Why it matters: The AI boom's constraint is shifting from GPUs to the HBM stacked on them, and whoever retains the memory-and-logic talent controls the supply that feeds every Nvidia accelerator.</description></item><item><title>Chinese DRAM maker CXMT surpasses Intel's market cap on a 500% debut</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1v7vdvg/chinese_chipmaker_cxmts_market_capitalization</link><guid isPermaLink="false">2026-07-27:hardware:https://www.reddit.com/r/LocalLLaMA/comments/1v7vdvg/chinese_chipmaker_cxmts_market_capitalization</guid><pubDate>Mon, 27 Jul 2026 07:00:00 +0000</pubDate><description>CXMT, mainland China's only integrated device manufacturer mass-producing general-purpose DRAM, surged nearly 500% on its first trading day to roughly RMB 3.28 trillion, the largest company by value on China's A-share market. That edges past Intel, which closed the prior day at about $465.6 billion (~RMB 3.15 trillion). The Hefei-based firm is central to China's push for domestic memory supply.

Why it matters: Memory is the bottleneck for AI accelerators; a well-capitalized domestic DRAM champion signals China intends to close the HBM and DRAM gap that export controls were meant to hold open.</description></item><item><title>Anthropic asks SK Hynix for supplies to build its own chips</title><link>https://fortune.com/2026/07/25/sk-chair-chey-tae-won-anthropic-chip-supplies-skhynix</link><guid isPermaLink="false">2026-07-26:hardware:https://fortune.com/2026/07/25/sk-chair-chey-tae-won-anthropic-chip-supplies-skhynix</guid><pubDate>Sun, 26 Jul 2026 07:00:00 +0000</pubDate><description>SK Group chair Chey Tae-won said Anthropic approached SK Hynix, one of the largest memory makers, for supplies to make its own semiconductors, speaking on stage alongside Dario Amodei at a San Francisco AI event. Chey called it remarkable for an AI developer to pursue its own silicon. The visit coincided with South Korea's president convening an AI summit, where Nvidia also announced partnerships with Naver and SK Group.

Why it matters: After committing to 2GW of AMD MI450s last week, Anthropic sniffing at custom silicon signals it wants leverage over both the Nvidia and AMD supply queues.</description></item><item><title>Etched raises $300M at $10.3B to build transformer-inference systems</title><link>https://techcrunch.com/2026/07/23/ai-chip-startup-etched-defies-skeptics-hits-10-3b-valuation-from-big-name-investors</link><guid isPermaLink="false">2026-07-24:hardware:https://techcrunch.com/2026/07/23/ai-chip-startup-etched-defies-skeptics-hits-10-3b-valuation-from-big-name-investors</guid><pubDate>Fri, 24 Jul 2026 07:00:00 +0000</pubDate><description>Etched closed a $300M Series C at a $10.3B valuation led by Sequoia, with a16z, SK Hynix, and Jane Street participating, doubling its December valuation in seven months. The company says it has already booked $1B in orders and is shipping full rack systems, not just chips, with a low-voltage prefill chip and a 'cluster-scale memory' interconnect for the decode phase. It pushes back on the perception that its silicon runs only specific LLMs, claiming support for MoE models and non-transformer designs like Mamba. Etched also opened an 80,000 sq ft, 10 MW facility in Milpitas, framing its pitch as 'run the world's inference.'

Why it matters: Inference-specialized silicon is graduating from thesis to booked revenue, and the more credible these alternatives get, the more pricing pressure Nvidia faces on the serving side.</description></item><item><title>Anthropic commits to 2GW of AMD MI450 GPUs; AMD invests up to $5B</title><link>https://the-decoder.com/anthropic-will-deploy-2-gigawatts-of-amd-gpus-for-claude-in-a-deal-worth-up-to-5-billion</link><guid isPermaLink="false">2026-07-23:hardware:https://the-decoder.com/anthropic-will-deploy-2-gigawatts-of-amd-gpus-for-claude-in-a-deal-worth-up-to-5-billion</guid><pubDate>Thu, 23 Jul 2026 07:00:00 +0000</pubDate><description>AMD will invest up to $5 billion in Anthropic, which in turn will deploy up to 2 gigawatts of Instinct MI450-series accelerators in Helios rack systems — MI455X GPUs paired with EPYC "Venice" CPUs, Pensando networking and ROCm — with the first gigawatt landing in H1 2027. AMD's stake is milestone-gated on deployment, echoing its 6GW OpenAI and 6GW Meta arrangements. A multi-year engineering program will use Claude to improve AMD's ROCm software, and AMD will run Claude internally across its dev teams.

Why it matters: It's another circular chip-lab financing loop, but it gives Anthropic a real second GPU source alongside Nvidia, Amazon Trainium and Google TPUs — and puts Claude to work hardening the weakest part of AMD's stack, its software.</description></item><item><title>Google reportedly bakes Gemini's architecture into 'Frozen v2' silicon</title><link>https://the-decoder.com/googles-frozen-v2-chip-reportedly-bakes-geminis-architecture-directly-into-silicon-for-efficiency-gains</link><guid isPermaLink="false">2026-07-21:hardware:https://the-decoder.com/googles-frozen-v2-chip-reportedly-bakes-geminis-architecture-directly-into-silicon-for-efficiency-gains</guid><pubDate>Tue, 21 Jul 2026 07:00:00 +0000</pubDate><description>Per The Information, Google is building a server chip internally called Frozen v2 that hardcodes parts of Gemini's model architecture (not its weights) directly into hardware, claiming 6-to-10x more tokens per watt than its current TPUs, with deployment targeted for 2028. New weights can still be loaded, so the chip survives model updates; an earlier Jeff Dean design that froze weights themselves was scrapped as too brittle. It is meant for internal inference only, and the report nudged Alphabet stock up about 3% ahead of earnings.

Why it matters: Inference margin is the new competitive front, and specializing silicon to a single architecture is the logical extreme of the efficiency race, at the cost of being locked to that architecture.</description></item><item><title>First loan backed by inference chips: $400M for SambaNova silicon</title><link>https://techcrunch.com/2026/07/17/why-the-first-gpu-financiers-are-turning-to-inference-chips-in-a-400-million-deal</link><guid isPermaLink="false">2026-07-18:hardware:https://techcrunch.com/2026/07/17/why-the-first-gpu-financiers-are-turning-to-inference-chips-in-a-400-million-deal</guid><pubDate>Sat, 18 Jul 2026 07:00:00 +0000</pubDate><description>AI inference cloud General Compute landed a $400M loan from Upper90, reportedly the first financing to use inference-specific chips as collateral — SambaNova's power-efficient SN50, which the startup claims runs 16x faster than GPU clouds. Upper90 pioneered GPU-backed lending with Crusoe in 2021; it's now betting the next wave is cheap inference for open models, outside Nvidia's ecosystem.

Why it matters: Capital markets are beginning to price non-Nvidia inference silicon as a financeable asset, a small crack in Nvidia's dominance and a signal that serving open models cheaply is becoming its own infrastructure category.</description></item><item><title>OpenAI's actual first device is a $230 light-up keyboard for Codex</title><link>https://techcrunch.com/2026/07/15/amid-hardware-legal-battle-openai-releases-a-230-keyboard-for-codex</link><guid isPermaLink="false">2026-07-16:hardware:https://techcrunch.com/2026/07/15/amid-hardware-legal-battle-openai-releases-a-230-keyboard-for-codex</guid><pubDate>Thu, 16 Jul 2026 07:00:00 +0000</pubDate><description>Days after reports of a screenless smart speaker, OpenAI's first branded hardware turned out to be the Codex Micro — a $230, 13-key mechanical keypad built with Work Louder and sold through OpenAI's merch store. Its RGB 'Agent Keys' show live status for up to six Codex threads (thinking, done, needs input, error), with a rotary dial to set an agent's reasoning level and a joystick to launch workflows. It's a limited run, ships via Bluetooth/USB-C around July 24, and is explicitly positioned as a novelty 'command center' for managing fleets of coding agents.

Why it matters: It's a gimmick, not the Jony Ive companion device — but the hardware design encodes a real workflow assumption: developers now juggle enough parallel agents that they need an ambient dashboard to see which one is stuck.</description></item><item><title>OpenAI's first device: a screenless speaker built to feel alive</title><link>https://the-decoder.com/openais-first-hardware-product-is-a-screenless-ai-speaker-designed-to-feel-alive</link><guid isPermaLink="false">2026-07-15:hardware:https://the-decoder.com/openais-first-hardware-product-is-a-screenless-ai-speaker-designed-to-feel-alive</guid><pubDate>Wed, 15 Jul 2026 07:00:00 +0000</pubDate><description>Bloomberg reports OpenAI's debut hardware product is a portable, screenless smart speaker pitched internally as a 'new type of home computer for the AI era.' It pairs a camera and sensors with the just-launched GPT-Live voice mode, and adds mechanical parts that physically move to make it seem lifelike. Unveiling is planned for later this year with a 2027 release; Apple's trade-secrets suit over hardware chief Tang Tan could delay it. It is reportedly the first of about five devices, including a phone replacement, a pendant, and home robotics.

Why it matters: A camera-equipped, always-listening, deliberately anthropomorphized device with access to your email is a very different threat model than a chatbot tab — and the same GPT-4o sycophancy that caused problems now ships with a motor.</description></item><item><title>Apple's OpenAI complaint: 400 poached staff, an auth bug, and prototypes at interviews</title><link>https://techcrunch.com/2026/07/13/the-wildest-allegations-in-apples-trade-secrets-lawsuit-against-openai</link><guid isPermaLink="false">2026-07-14:hardware:https://techcrunch.com/2026/07/13/the-wildest-allegations-in-apples-trade-secrets-lawsuit-against-openai</guid><pubDate>Tue, 14 Jul 2026 07:00:00 +0000</pubDate><description>Details from the 41-page filing sharpen the case first reported last week: Apple says 400+ ex-employees now work at OpenAI, that engineer Chang Liu exploited a 'rare' authentication bug to reach Apple's network weeks after leaving ('LOL, I found out I can access the [network storage]'), and that hardware chief Tang Tan had candidates bring CAD files and physical prototypes to interviews. Apple also alleges io used its confidential metal-finishing techniques by misleading a supplier. OpenAI: 'We have no interest in other companies' trade secrets.'

Why it matters: Strip the espionage framing and this is a talent-mobility fight — OpenAI is well-funded enough to ignore the Valley's no-poach norms, and discovery could set precedent for how AI labs recruit from incumbents.</description></item><item><title>SK Hynix's StreamDQ moves weight dequantization into HBM</title><link>https://semiengineering.com/near-memory-dequantization-architecture-in-custom-hbm-for-llm-inference-sk-hynix</link><guid isPermaLink="false">2026-07-14:hardware:https://semiengineering.com/near-memory-dequantization-architecture-in-custom-hbm-for-llm-inference-sk-hynix</guid><pubDate>Tue, 14 Jul 2026 07:00:00 +0000</pubDate><description>An SK Hynix paper proposes StreamDQ, a near-memory architecture that performs on-the-fly weight dequantization inside custom HBM for high-throughput, large-batch LLM inference. It reports up to 7.08x speedup and 90.23% lower energy on mixed-precision GEMM.

Why it matters: If dequantization happens in the memory subsystem rather than the GPU, quantized serving stops paying the bandwidth tax on every weight fetch — potentially a big lever for FP4 and mixed-precision inference at scale.</description></item><item><title>Caltech spinout claims a full 27B model running on an iPhone</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1uv54fv/compressed_version_of_qwen3627b_coming_from</link><guid isPermaLink="false">2026-07-13:hardware:https://www.reddit.com/r/LocalLLaMA/comments/1uv54fv/compressed_version_of_qwen3627b_coming_from</guid><pubDate>Mon, 13 Jul 2026 07:00:00 +0000</pubDate><description>PrismML, a Khosla-backed Caltech spinoff, says it compressed Alibaba's Qwen 3.6 27B from ~54GB to under 4GB and got it running on an iPhone 17 Pro, with open weights due next Tuesday. Crucially, it claims all 27B parameters stay active, versus Apple's own new on-device model that uses a sparse 20B architecture with only 1-4B active at a time. CEO Babak Hassibi says the technique shrinks models 'without hindering performance,' the usual claim that a benchmark will need to settle.

Why it matters: If the quality claim survives contact with real evals, a genuinely dense 27B on a phone changes the on-device ceiling from toy assistants to something that can run agents and code. Weights next week means the community can check the math fast.</description></item><item><title>$80 Tesla P100s ran silently noisy math in llama.cpp for years; a 3-line patch fixes it</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1uu6p9o/your_80_tesla_p100_has_been_doing_silently_noisy</link><guid isPermaLink="false">2026-07-12:hardware:https://www.reddit.com/r/LocalLLaMA/comments/1uu6p9o/your_80_tesla_p100_has_been_doing_silently_noisy</guid><pubDate>Sun, 12 Jul 2026 07:00:00 +0000</pubDate><description>A years-old llama.cpp CUDA bug forced the Pascal P100 (sm_60) down an fp16 math path that the GTX 10-series and P40 (sm_61) were long ago exempted from. Measured against fp32-reference logits on Qwen3.6-27B, the fix cut median KL divergence ~2300x (0.0023 to 0.000001) and lifted top-token agreement from 96.5% to 99.9% — with decode ~1.4% faster, since real workloads are GEMM/bandwidth-bound, not fp16-vector-bound. The patch simply extends the sm_61 exemption to sm_60; it's shipped in a turboquant fork because GGML bans AI-assisted contributions, and the bug was isolated by an agent loop running Fable 5.

Why it matters: P100s are ~$80 with 16GB HBM2 at 732 GB/s amid a DRAM crunch; a chunk of their reputation for 'worse' output was this bug, and the fix is measured only on sm_60 — not the all-GPUs panic some will read into it.</description></item><item><title>Apple sues OpenAI, alleging a 'coordinated campaign' to steal hardware secrets</title><link>https://the-decoder.com/apple-sues-openai-for-allegedly-running-a-coordinated-campaign-to-steal-trade-secrets-through-poached-employees</link><guid isPermaLink="false">2026-07-11:hardware:https://the-decoder.com/apple-sues-openai-for-allegedly-running-a-coordinated-campaign-to-steal-trade-secrets-through-poached-employees</guid><pubDate>Sat, 11 Jul 2026 07:00:00 +0000</pubDate><description>Apple filed suit in California federal court accusing OpenAI of a systematic effort to misappropriate trade secrets for its unreleased devices, naming hardware chief Tang Tan (ex-iPhone/Watch design lead) and former engineer Chang Liu. The complaint says 400+ ex-Apple staff now work at OpenAI, that Liu downloaded dozens of confidential hardware files on an Apple laptop he never returned, and that Tan told candidates to bring 'actual parts' to interviews. OpenAI denies any interest in others' trade secrets; io Products, the Jony Ive startup OpenAI bought for ~$6.5B, is also a defendant.

Why it matters: The 2024 ChatGPT-in-iOS partnership has fully collapsed into a talent-and-IP war, and the timing — with OpenAI's device slipping to 2027 and an IPO rumored — makes this more than a spat over departing engineers.</description></item><item><title>SK Hynix raises $26.5B in the largest-ever foreign US IPO</title><link>https://techcrunch.com/2026/07/10/sk-hynix-raises-26-5b-in-the-biggest-foreign-ipo-in-us-history-is-urged-to-build-new-us-fabs</link><guid isPermaLink="false">2026-07-11:hardware:https://techcrunch.com/2026/07/10/sk-hynix-raises-26-5b-in-the-biggest-foreign-ipo-in-us-history-is-urged-to-build-new-us-fabs</guid><pubDate>Sat, 11 Jul 2026 07:00:00 +0000</pubDate><description>The HBM memory maker sold 177.9M ADRs at $149 each on Nasdaq, raising $26.5B — topping Alibaba's 2014 record — with demand reportedly 7x oversubscribed and the stock opening 14% above price. Proceeds fund a new Korean fab, a packaging plant and EUV scanners to ease the AI-driven memory shortage. Commerce Secretary Lutnick is separately pressing SK Hynix and Samsung to build US fabs, while Micron pledged $250B in domestic manufacturing.

Why it matters: HBM is the real bottleneck behind every GPU order; a supplier flush with $26.5B and under US pressure to onshore is a signal about where inference capacity — and its cost — goes next.</description></item><item><title>ZML's LLMD promises peak inference across Nvidia, AMD, TPU, Apple and Intel</title><link>https://techcrunch.com/2026/07/08/hot-french-startup-zml-releases-free-product-to-speed-inference-across-lots-of-ai-chips</link><guid isPermaLink="false">2026-07-08:hardware:https://techcrunch.com/2026/07/08/hot-french-startup-zml-releases-free-product-to-speed-inference-across-lots-of-ai-chips</guid><pubDate>Wed, 08 Jul 2026 07:00:00 +0000</pubDate><description>Paris startup ZML, backed by Yann LeCun, launched LLMD, an inference server that runs open-source LLMs at (claimed) maximum speed across Nvidia, AMD, Google TPU, Apple Metal and Intel Arc silicon. The pitch is breaking vendor lock-in and letting shops mix cheaper or lower-power chips; ZML says it's co-designing silicon with European chipmakers like Axelera, SiPearl and VSORA. LLMD is free but not open source, launched to gather usage data. The 20-person team has raised ~$20M and enters a crowded field against vLLM, SGLang and Baseten.

Why it matters: A genuinely chip-agnostic inference layer would loosen Nvidia's grip and give infra teams real leverage on cost-per-token — if the cross-vendor performance claims survive independent benchmarks.</description></item><item><title>Qualcomm launches GenieX to run LLMs on Snapdragon Windows laptops</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1uo9z3c/qualcomm_launches_geniex_to_run_llms_on_their</link><guid isPermaLink="false">2026-07-06:hardware:https://www.reddit.com/r/LocalLLaMA/comments/1uo9z3c/qualcomm_launches_geniex_to_run_llms_on_their</guid><pubDate>Mon, 06 Jul 2026 07:00:00 +0000</pubDate><description>Qualcomm, late to the on-device SDK race, released GenieX for running LLMs across CPU, GPU, and NPU on its Windows laptops. Early hands-on reports: ~20 tok/s on Gemma 4 26B (A4B) with 0.5s to first token on GPU/NPU, and ~10 tok/s for Qwen 3.6 27B with MTP on GPU. Standard Q4_0 GGUFs reportedly run via llama.cpp on the CPU.

Why it matters: Usable NPU/GPU offload on mainstream Windows laptops widens the hardware base for local inference beyond Apple Silicon and discrete NVIDIA cards — if the tooling holds up in practice.</description></item><item><title>Anthropic in early talks with Samsung to build a custom AI chip</title><link>https://www.technology.org/2026/07/04/anthropic-samsung-custom-ai-chip</link><guid isPermaLink="false">2026-07-04:hardware:https://www.technology.org/2026/07/04/anthropic-samsung-custom-ai-chip</guid><pubDate>Sat, 04 Jul 2026 07:00:00 +0000</pubDate><description>The Information reports Anthropic is exploring a custom processor built on Samsung's 2nm process and advanced packaging, and has hired Clive Chan, an early member of OpenAI's silicon team. The project is very early: no design, testing, or defined function yet, and Anthropic insists Nvidia GPUs, Google TPUs, and AWS Trainium will remain central. Samsung, SK Hynix, and Micron were strategic investors in Anthropic's $65B Series H. The move follows OpenAI's Broadcom-built 'Jalapeño' inference chip unveiled last week.

Why it matters: Every frontier lab now wants leverage over Nvidia and its own performance-per-watt story; the question is whether Anthropic can ship silicon years behind Google and Amazon without derailing its rented-compute supply lines.</description></item><item><title>Anthropic in early talks with Samsung for a custom AI chip</title><link>https://www.theinformation.com/articles/anthropic-talks-samsung-manufacture-custom-ai-chip</link><guid isPermaLink="false">2026-07-03:hardware:https://www.theinformation.com/articles/anthropic-talks-samsung-manufacture-custom-ai-chip</guid><pubDate>Fri, 03 Jul 2026 07:00:00 +0000</pubDate><description>The Information reports Anthropic is discussing a custom accelerator with Samsung, though workloads, performance targets and process node are all undecided. Samsung offers its 4nm node and a data-center-tuned 2nm SF2P process entering production this year. Anthropic told press that AWS, Google and Nvidia silicon remains central to its strategy, and it has hired chip engineers including Clive Chan, an early member of Tesla's and OpenAI's silicon teams.

Why it matters: It follows OpenAI's Broadcom-built 'Jalapeño' inference chip by days: every major lab now wants custom silicon to escape Nvidia margins and control inference cost-per-watt. Whoever runs inference cheapest keeps more revenue.</description></item><item><title>Meituan's LongCat-2.0: 1.6T params trained entirely on domestic chips</title><link>https://www.scmp.com/tech/tech-trends/article/3358854/china-debuts-biggest-ai-model-trained-local-chips-meituan-releases-longcat-20</link><guid isPermaLink="false">2026-06-30:hardware:https://www.scmp.com/tech/tech-trends/article/3358854/china-debuts-biggest-ai-model-trained-local-chips-meituan-releases-longcat-20</guid><pubDate>Tue, 30 Jun 2026 07:00:00 +0000</pubDate><description>Meituan open-sourced LongCat-2.0, a 1.6-trillion-parameter model with a 1M-token context window, and claims it is the first trillion-parameter model to complete both pre-training and inference on a ~50,000-card domestic cluster of AI ASIC superpods. That goes a step beyond DeepSeek-V4-Pro, which Meituan says used home-grown chips only for inference. Pre-training is the far more compute-intensive phase, making the claim notable if it holds up.

Why it matters: If verified, it signals Chinese accelerators can handle frontier-scale training, not just inference — eroding one of the assumptions behind US export controls.</description></item><item><title>Chip geopolitics: Korea's $1T bet, Taiwan raids Super Micro</title><link>https://arstechnica.com/ai/2026/06/south-korea-to-spend-1t-on-more-memory-chip-production-and-humanoid-robots</link><guid isPermaLink="false">2026-06-30:hardware:https://arstechnica.com/ai/2026/06/south-korea-to-spend-1t-on-more-memory-chip-production-and-humanoid-robots</guid><pubDate>Tue, 30 Jun 2026 07:00:00 +0000</pubDate><description>South Korea committed $1 trillion across memory-chip production, AI data centers and humanoid robots, with President Lee calling semiconductors, physical AI and data centers the triple axis for a great leap forward. The same day, Taiwanese prosecutors raided Super Micro offices and partner firms over alleged smuggling of Nvidia AI chips into China; Super Micro's stock fell 8% and a co-founder was reportedly indicted.

Why it matters: The hardware supply chain is now an explicit instrument of state policy — both massive subsidies and criminal enforcement — and that volatility flows straight through to GPU and memory prices developers pay.</description></item><item><title>Samsung and SK Hynix commit ~$518B to new chip hub for AI demand</title><link>https://abcnews.com/Technology/wireStory/south-korean-tech-giants-build-518-billion-chipmaking-134300835</link><guid isPermaLink="false">2026-06-29:hardware:https://abcnews.com/Technology/wireStory/south-korean-tech-giants-build-518-billion-chipmaking-134300835</guid><pubDate>Mon, 29 Jun 2026 07:00:00 +0000</pubDate><description>Samsung and SK Hynix, backed by the South Korean government, will invest a combined 800 trillion won (~$518B) in a new chipmaking hub in the country's southwest, with each building two fabs; The Decoder puts the total program nearer $590B including packaging and next-gen chip spending. The two firms control roughly 80% of the high-bandwidth memory market AI workloads depend on. Jefferies expects memory prices to rise 40-50% in Q3 2026 and another 30-40% in Q4, with relief unlikely before 2028.

Why it matters: HBM and DRAM price spikes are already pushing up hardware costs (Apple has hiked Mac prices), so anyone budgeting GPU or local-inference builds should expect memory to stay expensive into 2027.</description></item><item><title>Everyone wants off Nvidia: OpenAI's Jalapeño joins the custom-silicon rush</title><link>https://techcrunch.com/video/why-everyone-from-openai-to-spacex-is-building-their-own-chips-and-turning-up-the-heat-on-nvidia</link><guid isPermaLink="false">2026-06-27:hardware:https://techcrunch.com/video/why-everyone-from-openai-to-spacex-is-building-their-own-chips-and-turning-up-the-heat-on-nvidia</guid><pubDate>Sat, 27 Jun 2026 07:00:00 +0000</pubDate><description>OpenAI detailed Jalapeño, a custom inference chip built with Broadcom, joining Google, Apple, and SpaceX in building their way out of single-supplier risk. The framing is hedge, not clean break — more control and hardware tuned to specific workloads, echoing Apple's gains from dropping Intel. The same discussion noted Groq raising $650M after Nvidia poached its top talent.

Why it matters: Custom inference silicon from the largest API providers could reshape pricing and availability downstream. If Jalapeño lands, it's another lever OpenAI gains over the cost curve that determines what you pay per token.</description></item><item><title>PyTorch's TokenSpeed-kernel makes multi-silicon inference a registry problem</title><link>https://pytorch.org/blog/lightseek-tokenspeed-kernel</link><guid isPermaLink="false">2026-06-26:hardware:https://pytorch.org/blog/lightseek-tokenspeed-kernel</guid><pubDate>Fri, 26 Jun 2026 07:00:00 +0000</pubDate><description>A PyTorch blog details TokenSpeed-kernel, a standalone kernel subsystem that decouples the inference runtime from hardware-specific code via a public API (mha_prefill, moe_apply, etc.) plus a registry-and-selector that dispatches to platform kernels. Using GPT-OSS 120B on AMD MI355X (CDNA4) as the test case, Gluon-backed attention and MoE kernels delivered 1.6–3.6x end-to-end throughput over the portable Triton path, with the AMD kernels published separately as tokenspeed-kernel-amd and already adopted by vLLM. NVIDIA Blackwell paths sit behind the same API via FlashInfer/TensorRT-LLM wrappers.

Why it matters: Backend selection leaking into model code is a real maintenance tax as GPU vendors, quant formats, and architectures multiply. A clean kernel boundary that vLLM can borrow is how AMD stays a first-class inference target rather than a perpetual afterthought.</description></item><item><title>OpenAI and Broadcom tape out 'Jalapeño,' a custom LLM inference chip</title><link>https://arstechnica.com/gadgets/2026/06/openai-and-broadcom-announce-chip-designed-for-llm-inference-at-scale</link><guid isPermaLink="false">2026-06-25:hardware:https://arstechnica.com/gadgets/2026/06/openai-and-broadcom-announce-chip-designed-for-llm-inference-at-scale</guid><pubDate>Thu, 25 Jun 2026 07:00:00 +0000</pubDate><description>OpenAI unveiled Jalapeño, its first custom accelerator (an 'Intelligence Processor') built with Broadcom specifically for LLM inference, with OpenAI doing chip design and Broadcom contributing silicon and Tomahawk networking. OpenAI claims design-to-tape-out took nine months — partly accelerated by its own models — and 'substantially better' performance per watt, though these are self-reported numbers with no technical report yet. Engineering samples are already running GPT-5.3-Codex-Spark in the lab; large-scale deployment is planned for late 2026 at gigawatt scale, with Microsoft reportedly committed to buying 40% of the first run. Community reverse-engineering pegs it as TPU-like, roughly 216GB HBM3E and ~10 PFLOPS FP4.

Why it matters: If the perf-per-watt claims hold, OpenAI gains leverage over inference economics and its Nvidia dependence — but until an independent technical report lands, treat the numbers as marketing.</description></item><item><title>Qualcomm enters the data center with Dragonfly C1000 and buys Modular for ~$4B</title><link>https://the-decoder.com/qualcomm-enters-the-data-center-market-with-its-own-processor</link><guid isPermaLink="false">2026-06-25:hardware:https://the-decoder.com/qualcomm-enters-the-data-center-market-with-its-own-processor</guid><pubDate>Thu, 25 Jun 2026 07:00:00 +0000</pubDate><description>Qualcomm announced the Dragonfly C1000, a data-center processor optimized for AI agents and low power, with Meta planning to deploy it starting 2028. Alongside it, Qualcomm is acquiring Chris Lattner's Modular — maker of the cross-architecture Mojo/inference stack — for roughly $4 billion, with Modular saying Mojo open-sourcing stays on track. Qualcomm nearly doubled its non-smartphone revenue forecast to $40B by 2029 (targeting $15B from data centers); the stock jumped 15% after hours.

Why it matters: The Modular buy gives Qualcomm a serious CUDA-alternative software story to pair with its silicon — another front in the slow erosion of Nvidia's lock-in.</description></item><item><title>Seven Chinese vendors are now shipping H100/H200-class accelerators</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1udkxde/7_chinese_companies_are_already_shipping</link><guid isPermaLink="false">2026-06-24:hardware:https://www.reddit.com/r/LocalLLaMA/comments/1udkxde/7_chinese_companies_are_already_shipping</guid><pubDate>Wed, 24 Jun 2026 07:00:00 +0000</pubDate><description>A widely-shared LocalLLaMA writeup maps at least seven Chinese AI-chip makers shipping today: 'three dragons' (Huawei Ascend, Alibaba T-Head, Baidu Kunlunxin) and 'four snakes' that mostly IPO'd in the last six months (MetaX, Moore Threads, Biren, Iluvatar CoreX). Current parts land around H100, next-gen targets H200, and production is shifting from TSMC to SMIC. The post cites a CHITEX talk for many specifics and flags vendor/analyst figures as unverified. NVIDIA's China GPU share reportedly fell from 95% to 55% in two years. Separately, a Chinese supercomputer reclaimed the world's-fastest spot for the first time since 2017.

Why it matters: Chinese open-weight models (Qwen, DeepSeek, GLM) are increasingly co-designed with domestic silicon, with its own form factor, interconnect, and HBM. If you run open weights, the hardware you target in two years may not be NVIDIA.</description></item><item><title>Reflection rents $6.3B of GB300s from SpaceX, the third neocloud deal</title><link>https://techcrunch.com/2026/06/22/spacex-inks-compute-deal-with-reflection-ai-an-open-source-ai-lab</link><guid isPermaLink="false">2026-06-23:hardware:https://techcrunch.com/2026/06/22/spacex-inks-compute-deal-with-reflection-ai-an-open-source-ai-lab</guid><pubDate>Tue, 23 Jun 2026 07:00:00 +0000</pubDate><description>Open-weight lab Reflection AI will pay SpaceX $150M/month from July 2026 through 2029 for immediate access to Nvidia GB300 chips at the Colossus 2 data center near Memphis — a deal worth up to $6.3B, with a 90-day exit clause. It is smaller than SpaceX's Anthropic ($1.25B/month) and Google ($920M/month) contracts. Tallied together, SpaceX's GPU rentals annualize to roughly $28B/year at implied Blackwell pricing above $10/hour, about twice CoreWeave's current revenue.

Why it matters: SpaceX has quietly become a major 'neocloud,' and GPU brokerage is emerging as a strategic layer between model builders and hardware supply — with Reflection pitching open weights as the hedge against closed-model access being revoked.</description></item></channel></rss>
