Apple's new Siri runs on Google's Gemini

Apple finally shipped its rebuilt Siri — and the headline detail for developers is that it leans on Google's Gemini rather than Apple's own frontier models. The frontier labs' slowdown pitch hardened into a concrete third-party-evaluator standard even as Cohere and Hugging Face branded it a cartel. Meanwhile local agents kept marching onto laptops, a fresh crop of open weights landed, and a "world's cheapest" inference provider imploded on air.

Apple ships its rebuilt Siri, built on Google's Gemini

Apple released iOS 27, macOS 27 Golden Gate and the rest of the 2026 OS lineup, with a large-model Siri overhaul as the flagship feature. Per The Decoder and TechCrunch, the new 'Siri AI' is built on Google's Gemini models running through Private Cloud Compute, while on-device work uses Apple's own AFM 3 models — AFM 3 Core (3B) and a sparse AFM 3 Core Advanced (20B, activating 1–4B per request). Siri can read on-screen content, pull context from messages and photos, and trigger system-wide app actions across first- and third-party apps. It launches in English only and is withheld from the EU and China for now.

Why it matters: Apple conceding the foundation-model layer to Google is the story: the company that pitched on-device privacy now routes its assistant through a rival's cloud model. Developers get a system-wide app-actions surface worth targeting once Siri AI stabilizes.

Slowdown pitch hardens into an evaluator standard — and draws 'cartel' fire

The frontier-labs pacing debate moved from essays to mechanisms. The AI Evaluator Forum published AEF-1, a baseline for independent third-party evaluations covering access, conflicts of interest and recusal, and Anthropic said it will unilaterally give embedded evaluators like METR employee-level access — 'desks in our offices, access badges, and company laptops.' The pushback was fierce: Cohere CEO Aidan Gomez called the antitrust-exemption plan 'a cartel by any other name,' a Hugging Face engineer called it 'bizarre nonsense,' and David Sacks said tying a slowdown to a preferred regulatory framework 'will look like blackmail.' Trump again dismissed AI-takeover warnings as a hoax.

Why it matters: The concrete artifact here is AEF-1 and embedded-evaluator access — a governance template that could bind anyone building at the frontier. The unresolved question is whether evaluators funded by the labs they audit can be independent.

Perplexity puts a local agent on Windows RTX PCs

Perplexity's Portable Computer — a local version of its agentic Computer product — is now available in the Windows app on NVIDIA GeForce RTX and RTX PRO systems with 24GB+ VRAM, extending earlier DGX Spark and Linux support. It runs a Qwen 3.8 27B model post-trained for the agent and keeps sensitive files on-device, with locally completed work not consuming cloud credits; the agent asks permission before escalating a task to cloud models. Connectors cover Outlook, OneDrive, Google Drive, Gmail, Slack and GitHub.

Why it matters: This is a concrete data point on where local agents are usable today: a 27B model on a consumer GPU handling multi-step file and code chores, with cloud escalation as an explicit opt-in rather than the default.

IFM's K2 Horizon open weights land, with a KV-cache catch

The full K2 Horizon lineup from IFM appeared on Artificial Analysis and Hugging Face, with r/LocalLLaMA users reporting the 3.7B and 7B models as unusually strong for their size — one thread claims the 7B ranks between Qwen 3.6 27B and 35B-A3B and that IFM open-sourced every training step. A detailed community teardown by crusaderky cautions the models carry a 'god-awful KV cache design': the 3.7B and 7B each need ~5 GiB just for 128k context, so a params-based 'best in class' read shifts sharply once you plot RAM instead.

Why it matters: Small open models keep creeping up the intelligence-per-byte curve, but the context-memory footprint — not parameter count — is what decides whether they fit your VRAM. Benchmarks alone will mislead here.

Sakana trains 1,000-layer nets without backpropagation

Sakana AI published PC-ALM (Augmented Lagrangian Predictive Coding), a backprop alternative that trains deep networks using only layer-local dynamics. By adding per-layer Lagrange multipliers to standard predictive coding, the method recovers exact backprop credit signals in linear networks and, Sakana says, trains residual MLPs up to 1,000 layers while nearly matching backprop, improving on plain predictive coding across MNIST, CIFAR-10 and Tiny ImageNet. The stated motivation is biological plausibility and energy-efficient training on neuromorphic hardware.

Why it matters: Local-only credit assignment at 1,000 layers is a real milestone for the predictive-coding line of work, and a reminder that the field is still probing alternatives to the backprop-on-GPU orthodoxy — though the results remain on small tasks.

Anthropic's $2T IPO stays on track, headed for Nasdaq

Sources tell Axios that Anthropic still plans to go public in 2026 despite the AI-safety uproar, and Business Insider reports it has chosen the Nasdaq. Anthropic told investors it will post a second straight profitable quarter — on an adjusted metric that excludes stock-based compensation — with gross margins above 80% before revenue-sharing and training costs, per the FT. Quarterly revenue reportedly hit $11.5 billion (14x year over year) for a $65 billion annualized run rate at the end of July, against a possible $2 trillion-plus valuation. Sam Altman, meanwhile, said OpenAI won't IPO this year.

Why it matters: The pacing rhetoric hasn't dented the capital-markets plan — and a cynical reading is that 'slowing down' could cut compute spend and improve the very financials Anthropic is taking to public investors.

'Cheapest inference in the world' was an OpenRouter wrapper

According to an exposé on kendell.dev widely circulated on r/LocalLLaMA, CrofAI (crof.ai / nahcrof.com) — which billed itself as the world's cheapest inference provider and dismissed rivals' pricing as 'skill issues' — was allegedly a thin OpenRouter wrapper that silently routed requests to cheaper, weaker models at up to a 20x markup, with a 'greg' house model family that mapped to GLM and Qwen checkpoints. The write-up documents physically implausible hardware claims and five failed attempts to hide the OpenRouter fingerprints. After the report, the operator announced a shutdown, floated a fake 'new team' takeover, then wiped the site, Twitter account and subreddit within hours.

Why it matters: A cautionary tale for anyone chasing rock-bottom token prices: if a provider's economics look impossible, verify what model is actually answering. Fingerprinting and independent benchmarks beat marketing copy.

Browse previous days →