Xiaomi crashes the open-weights frontier
A phone maker, not one of the established Chinese labs, shipped the day's top open-weights model — and did it with a reinforcement-learning run that reportedly cost single-digit millions. Elsewhere the frontier gap showed its teeth: xAI's Grok 4.7 landed cheap but mid-pack, and Amazon slammed the door on Meta's shopping agent. OpenAI stood up a math advisory board it can overrule, and Cloudflare made Python a first-class edge language.
Xiaomi's MiMo-V2.6 debuts as the top open-weights model, trained on a cheap RL run
Xiaomi released MiMo-V2.6, a natively omnimodal open-weights family under an MIT license. The Pro model carries 1.02T total and 42B active parameters and, per Artificial Analysis, debuts as the top open-weights model on its Intelligence Index at 46, priced at $0.435 per million input and $0.87 per million output tokens. Xiaomi published weights, a technical report, and its RL training code and environments (with roughly 7,000 tasks promised); a widely cited figure puts the final RL run at about 130 hours, 75B tokens and $2.6M. The team also shipped a MiMo-V2.6-Distill-Qwen-9B.
Why it matters: If frontier-adjacent results really come out of a few-million-dollar RL run plus open environments, post-training rather than pretraining scale becomes the cheap lever — and a phone maker just out-shipped the six established Chinese AI labs on it.
- [AINews] Xiaomi MiMo-V2.6-Pro 1T-A42B: the new top Open Weights model, trained for $3M (Latent Space)
- MiMo-V2.6 distilled themselves into Qwen 9B (r/LocalLLaMA)
- Mimo v2.6-Flash-RL vs open-weight models (r/LocalLLaMA)
Grok 4.7 ships cheap, benchmarks land mid-pack and well behind on agentic coding
xAI launched Grok 4.7 at $2 per million input and $6 per million output tokens, built on a larger base model with a longer reinforcement-learning run and better self-verification. On the independent Artificial Analysis Intelligence Index it scores 46, mid-pack, against 53 each for Claude Fable 5.1 and GPT-6. On Terminal-Bench 4.0 it manages just 26% versus 60% for GPT-6 Astra and 55% for Fable 5.1 — even DeepSeek V4.1 Flash edges it at 27%. xAI claims an all-new safeguard stack, topping LatchBio's biosafety benchmark at 62.4% and its own HackerBench cyber test.
Why it matters: The pricing sits at Chinese-model levels, and the benchmarks suggest that is the point: agentic coding is precisely where Grok's gap to the closed frontier is widest.
Amazon blocks Meta's Muse agent from shopping on amazon.com
Amazon has cut off Meta's new Muse assistant, returning an error that continued access "by an unauthorized AI agent" violates its terms of use. Amazon says Muse browsed without permission, does not identify itself as an AI, and appears to store customer data, calling it a security and privacy risk; Meta says Muse has no visibility into passwords or payment methods. It extends Amazon's pattern of barring agentic shoppers — it sued Perplexity's Comet last year, a ban an appeals court overturned in August — even as Meta remains a billion-dollar Amazon cloud-chip customer. Muse launched September 8 and topped the US App Store within a week.
Why it matters: Agentic commerce keeps hitting the same wall: the marketplace, not the model vendor, has to clean up a hallucinated order, so the big retailers are refusing agents at the door regardless of partnership ties.
- Meta's AI agent has been blocked from using Amazon.com (TechCrunch)
- Amazon bars Meta's AI agent from online shopping (Morning Brew)
- Amazon blocks Meta's AI agent Muse from online shopping (The Decoder)
Cloudflare makes Python Workers generally available, databases and LLM SDKs included
After a two-year preview, Cloudflare made Python Workers GA, with Python now a first-class Workers language via Pyodide compiled to WebAssembly. You can run FastAPI, Django and Flask through built-in ASGI/WSGI connectors, reach PostgreSQL and MySQL via Hyperdrive (Cloudflare implemented socket syscalls over its connect API), and run openai, langchain and mcp natively after upstream fixes route their HTTP through the JavaScript fetch API. The effort produced PEP 783, standardizing a PyEmscripten platform so any Python package can build wheels for the browser/Wasm runtimes.
Why it matters: Edge Python that talks to real relational databases and LLM SDKs without JavaScript glue makes Workers a plausible home for serverless AI backends, and the Pyodide upstreaming benefits the whole Python-on-WebAssembly ecosystem.
- Python Workers are now generally available (Cloudflare Blog)
- Cloudflare Python Workers are now generally available (Simon Willison)
OpenAI forms a math advisory group it can't be overruled by, claims 100+ solved problems
OpenAI announced an independent Advisory Group on Mathematics and AI, hosted at the Institute for Advanced Study in Princeton, and alongside it claimed an internal model has resolved more than 100 open math problems, following its earlier Navier-Stokes solution. The nine-member group can assess and coordinate the release of results but, per OpenAI and the IAS, explicitly cannot slow or redirect the company's research. Only one member, Camillo De Lellis, signed a recent open letter from 25 Fields Medalists objecting to the labs' pace. The 100-problem claim remains among the least independently evaluated results in circulation.
Why it matters: A body that advises but cannot say "stop" looks more like release management than oversight — and a sweeping unverified problem count is exactly the sort of claim mathematicians are asking labs to substantiate.
Alibaba unveils Qwen 4 and a new AI chip at Apsara, teases a multi-trillion-parameter model
At its Apsara conference Alibaba announced Qwen 4 and detailed a new in-house AI chip, while laying out plans to scale its flagship model. Reports indicate a planned model in the 5-trillion to 10-trillion-parameter range. Concrete specifications for both Qwen 4 and the accelerator remain thin, and much of the parameter detail comes from secondhand summaries rather than Alibaba's own materials.
Why it matters: Pairing a custom accelerator with multi-trillion-parameter ambitions is Alibaba's bid to blunt its Nvidia dependence and defend Qwen's lead among open Chinese models.
Also worth a look
- Jev introduces a new shape of LLM - System One, aka Decision Models (Simon Willison)
- Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI (Latent Space)
- TinyTorch: Don't Just Import PyTorch. Build It. (PyTorch)
- Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community (Hugging Face)
- M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents (MacStories)
- Can gzip be a language model? (nathan.rs)
- Anthropic, OpenAI call for Australia to relax ban on training of AI models (Reuters)
- New York tightens AI regulations as RAISE Act reporting begins in November (13WHAM)