Qwen 3.8 lands as Kimi buckles
China's open-weight labs dominated the day from both ends: Alibaba dropped a 2.4-trillion-parameter Qwen 3.8 that it claims trails only Fable 5, while Moonshot's Kimi K3 buckled under its own launch demand and froze new signups. Elsewhere, Hugging Face admitted it fought off an AI-driven breach with a Chinese open model after US frontier APIs refused the job, and fresh research showed LLMs invent hiring biases no human taught them.
Alibaba ships Qwen 3.8, a 2.4T open-weight model it rates second only to Fable 5
Qwen 3.8 is a 2.4-trillion-parameter model and the team's first multimodal release above 1T params, handling images, video and documents. It landed as a paid preview via Alibaba's Token Plan, Qoder and QoderWork at 10 percent of standard price, with open weights promised 'soon' and no independent benchmarks yet. Early hands-on reports praise its coding but flag frequent thinking loops, and the timing directly targets Kimi K3's momentum.
Why it matters: A genuinely open 2.4T multimodal model at preview pricing would reset the price/capability floor for self-hostable coding, but 'second only to Fable 5' is a vendor claim with zero public numbers and visible loop bugs — treat it as a preview, not a benchmark.
Kimi K3 freezes new subscriptions 48 hours in as demand outruns GPUs
Moonshot paused new Kimi K3 consumer subscriptions after requests 'pushed close to the limits of our current capacity,' prioritizing existing paid users and splitting plans into a general 'Kimi Membership' and a separate 'Kimi Code Membership' to ration compute. Reuters reports the crunch coincides with a fresh $2B raise at a $30B valuation and preparations for a Hong Kong IPO. Analysts note K3's 2.8T size and agentic, multi-call workloads make it expensive to serve — and impractical for most to self-host despite the open weights.
Why it matters: So much for open weights cutting compute needs: the largest open model to date is capacity-constrained days after launch, a reminder that 'open' doesn't mean 'runnable' at 2.8T and that hosted access, not the download, is where the business lives.
Hugging Face fought an AI-driven breach with a Chinese open model after US APIs refused
Hugging Face disclosed a July breach in which an autonomous AI agent system chained two code-execution paths in its dataset processing, escalated to node-level access, harvested cloud credentials and moved laterally across clusters via short-lived sandboxes. When responders fed the 17,000+ attack logs to commercial frontier APIs, safety guardrails blocked the analysis — so they ran forensics on Z.ai's open-weight GLM 5.2 on their own infrastructure, which also kept attacker data in-house. The company advises rotating access tokens and pre-vetting a self-hostable model before an incident.
Why it matters: This is the concrete case open-weight advocates have been waiting for: refusal classifiers tuned to trip on anything that looks offensive also lock out the blue team, making a capable local model an incident-response requirement, not a preference.
LLMs invent hiring biases no human taught them, ICML study finds
Princeton and University of Chicago researchers ran ChatGPT, Claude, Gemini and others through a 40-round simulated hiring game where all candidates were equally likely to succeed. The models rapidly segregated four fictional ethnic groups into job niches from a handful of early outcomes, scoring ~65% higher on a segregation scale than human participants (o3 hit 1.83, near the 2.0 max). Telling models to be fair barely helped; offering a diversity bonus, or supplying relevant personal detail, did.
Why it matters: As vendors race to ship agents with persistent memory, this shows personalization is also a bias-accumulation surface — a résumé-screening agent can over-index on its own past outcomes and manufacture discrimination from noise, with no training-data smoking gun to audit.
- AI is more likely than humans to form biases when hiring (MIT Technology Review)
Musk v. Altman exposes 2022 email: OpenAI's open-source plan was to freeze out rivals
A newly surfaced October 2022 email from Sam Altman to OpenAI's board, exposed in the Musk v. Altman litigation, proposes releasing a locally-runnable GPT-3-class model — explicitly to 'discourage others from releasing similarly-powerful models' and make it 'harder for new efforts to get funded.' Simon Willison flagged the quote as a candid window into how open releases were pitched internally as a competitive moat rather than a gift.
Why it matters: Against a backdrop of OpenAI execs now warning about Chinese open weights, the 2022 framing lands differently: openness was a strategic lever the whole time, useful context for reading today's 'open-source is dangerous' arguments.
- Quoting Sam Altman (Simon Willison)
MiniCPM goes embodied with open-source VLA and tracking models
OpenBMB open-sourced MiniCPM-Robot, its first embodied-AI series: MiniCPM-RobotManip, a 1.5B general-purpose vision-language-action model for robotic manipulation, and MiniCPM-RobotTrack, a 0.5B model for real-world target tracking. The release ships alongside PhyAI, an inference framework built for embodied models, with weights on Hugging Face.
Why it matters: Sub-2B open VLA models that target real robot hardware push embodied AI toward hobbyist and edge budgets, and give developers a concrete open baseline to fine-tune against instead of closed robotics stacks.
OpenAI regains secondary-market bid on GPT-5.6 and Codex, but Anthropic still leads 5-to-2
Secondary-market traders report a 'resurgence' in demand for OpenAI shares after the GPT-5.6 Sol/Terra/Luna launches and Codex plus ChatGPT Work hitting 9 million active users. OpenAI is valued around $933B (up ~20% in three months) versus Anthropic's ~$1.2T, with buyers still favoring Anthropic roughly five-to-two. Independent benchmarks place GPT-5.6 Sol near the top but below Claude's Mythos and Fable.
Why it matters: Private-market sentiment is a noisy proxy, but the Codex/ChatGPT Work usage figure is the concrete signal — evidence that agentic coding is spreading past the developer core into broader knowledge work.
Also worth a look
- Bezos backs CuspAI as startup teams up with Nvidia to hunt for chipmaking materials (CNBC)
- OpenAI & ReliaQuest: Partnership for Agentic Cybersecurity (AI Magazine)
- Fractale-350M-base: memory as trained behaviour instead of long context, a fully open research release (r/LocalLLaMA)
- [Paper] xHC: Expanded Hyper-Connections — scaling residual streams beyond N=4 (r/LocalLLaMA)
- [Paper] ATSInfer: Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices (r/LocalLLaMA)
- Silicon Valley is embracing armed robots. Washington is making room. (The Washington Post)
- BeeLlama.cpp v0.4.0: KVarN and KV precision tail for tighter KV-cache quantization (r/LocalLLaMA)
- A knowledge-enhanced domain-aware LLM agent for atrial fibrillation management (Nature)