Microsoft's own exec called scraping theft

Unsealed filings in the New York Times copyright case put Microsoft's and OpenAI's private views of news scraping on the record, in their own words. Elsewhere the hardware and safety threads kept moving: Huawei pulled its next Ascend chip forward to Q1 2027, and Google DeepMind spun up an institute to argue about AGI in public. On the tooling side, Anthropic pushed Claude Code further toward hands-off parallel agents.

Unsealed NYT filings quote Microsoft calling AI scraping 'the largest theft of labor in human history'

A newly unsealed summary-judgment brief in the New York Times' three-year-old suit against OpenAI and Microsoft surfaces internal documents the companies had kept confidential. In a January 2023 memo, Microsoft applied-science director Brent Hecht called the training practice 'an astonishing theft of unprecedented proportions' and 'the largest theft of labor in human history.' The filing cites specifics: OpenAI mid-training datasets allegedly holding 91,692 copies of NYT, Daily News and CIR works; a Common Crawl-derived set with over 2 million nytimes.com documents; and Copilot cutting click-through to the NYT domain by as much as 93% versus Bing search. Many quotes come from the plaintiffs' own brief, stripped of original context; the underlying exhibits remain sealed, and OpenAI and Microsoft did not comment.

Why it matters: The admissions cut directly at the fair-use defense the industry is leaning on, particularly the market-harm prong, and the Trump administration filed in OpenAI's defense earlier this month. If they survive context, they reshape the leverage in every training-data suit.

Huawei pulls Ascend 960DT forward to Q1 2027 — but the SuperPoD shrinks

At Huawei Connect, acting chairman David Wang said the next-generation Ascend 960DT AI chip is now expected in Q1 2027, moved up from a previously planned Q3, with claimed doubled performance. Huawei is pitching its Peerium Computing Architecture and UnifiedBus interconnect to lash chips into one machine; it says an Atlas 950 SuperCluster can link up to 256,000 accelerator cards. Analyst Rui Ma flagged a catch: the SuperPoD announced this week tops out at 4,096 chips, far below the 15,488-chip Atlas 960 SuperPoD Huawei had earlier described. So the chip arrives sooner, but the system around it is smaller than promised.

Why it matters: Huawei is the clearest test of whether export controls actually slow China's AI hardware. A pulled-in timeline days before the Trump–Xi meeting is as much signaling as engineering — and the shrunken cluster is the detail to watch.

Google DeepMind launches an institute, and Hassabis floats a US frontier-standards body

Google and Google DeepMind stood up the DeepMind Institute, with Shane Legg, James Manyika and Demis Hassabis as directors, publishing an opening set of four essays meant to air disagreement about AGI. Hassabis proposes a US-led frontier standards body: developers would first submit models voluntarily for review up to 30 days before release, with eventual mandatory, 'held-out' undisclosed evaluations to stop labs teaching to the test, and a framework that could be 'ratcheted up' to a coordinated slowdown. A separate essay by Rohin Shah and Anca Dragan argues the shrinking window to read a model's reasoning is not inevitable, and floats capping 'opaque serial depth.'

Why it matters: This turns the week's abstract slowdown talk into concrete institutional proposals — pre-release review, held-out evals, transparency limits — the shape any actual regulation would take. Coming from DeepMind, it's a competing blueprint to Anthropic's and OpenAI's.

Anthropic rebuilds Claude Code Projects around parallel cloud agents

Anthropic reworked the Projects feature in Claude Code so a user states a goal and a coordinator splits it across parallel 'threads,' each running as its own cloud session that can open pull requests and run tests. Progress is trackable per thread or in the main chat, including on mobile, and Claude builds shared memory across threads over time. The beta is limited to select Pro and Max subscribers using cloud sessions; Team and Enterprise access and local execution are slated to follow. It lands shortly after Anthropic made autopilot mode the default in Claude Code.

Why it matters: Fan-out-and-verify is becoming the default shape of agentic coding tools. It also, as The Decoder notes, shifts control over how many tokens get burned from the developer to the vendor — convenient timing for a company heading to IPO.

Crates security team warns of a social-engineering campaign against prominent Rust maintainers

Adam Harvey and the crates.io security team warn of an ongoing campaign targeting rust-lang members and owners of popular crates, aiming to compromise devices and accounts to publish malware. The lure is a video call framed around a job, project or contract, then used to get the target to install something (a supposedly missing audio codec) or run a command pasted onto their clipboard. The team says the same trick was used last month in a successful supply-chain attack on the array_ref crate, among others. Simon Willison's suggested defense: dependency cooldowns, holding off a few days before upgrading to new releases.

Why it matters: Every dependency graph is also a graph of humans with publish rights, and they're now being hunted directly. Adding a cooldown window before pulling fresh releases is a cheap, immediate mitigation any team can adopt today.

Baseten's Base Labs teams with Hugging Face and Goodfire on open-weight safety infrastructure

Baseten launched a safety-infrastructure standard alongside its Base Labs research arm, partnering with Hugging Face and Goodfire AI to build evaluation and monitoring tooling for open-weight models. The pitch is that safety should be trained into open models and enforced by whoever serves them, rather than bolted on afterward. The backdrop is abliteration — stripping safeguards from released weights — with Hugging Face already hosting over 6,000 abliterated models. Technical details of the partnership are not yet disclosed; Goodfire, an interpretability shop, is the likely candidate for the 'built-in' monitoring piece. Baseten raised a $1.5B Series F in June at a $13B valuation.

Why it matters: It's a rare attempt to make 'open weights' and 'safe' compatible at the serving layer, where inference providers actually sit. Whether it becomes a real standard or a marketing frame depends on the technical spec they haven't published yet.

Cactus's Needle 3 is a 121M on-device model that only makes function calls

In a detailed r/LocalLLaMA post, Henry from Cactus Compute introduced Needle 3, a 121M-parameter on-device 'automation' model that refuses to chat: every turn is a tool call, structured extraction or embedding, and a request no declared tool can serve returns an empty list rather than a guess. Arguments are emitted under a byte-level grammar compiled from the schema, so JSON always parses and enums can't escape their set. He claims 86.0 on Mobile Actions through the shipped 2-bit binary, against 82.4 for LFM2.5 1.2B and 88.4 for cloud DeepSeek V4 Flash, with 8–29MB binaries running on plain CPU up to 4k tokens/sec on a Raspberry Pi 5. One set of weights is sliceable to any depth from 2 to 20 layers. All figures are the vendor's own, self-reported.

Why it matters: Constrained-decoding tool-callers small enough to run air-gapped on a watch are a distinct bet from shrinking chat models, and the grounding rules (omit rather than invent) are exactly what agent plumbing wants. Treat the benchmark numbers as claims until someone reproduces them.

'Infinite-Parameter LLMs' propose writing live interaction into the weights

A new arXiv paper pitches an 'Infinite-Parameter LLM' that learns from run-time data by generating weights rather than storing them. Taking inspiration from Mixture-of-Experts, a compact hypernetwork turns data supplied during a session into a low-rank modulation of a shared base network, so feed-forward weights are compiled from live input instead of read from a fixed bank. Where prior weight generators read context once and freeze, the authors carry a Bayesian belief over the generator's latent code and update it online, re-deriving the effective weights as the session proceeds. The stored footprint stays fixed; the paper specifies an evaluation protocol pitting the approach against in-context learning and retrieval.

Why it matters: It's a concrete alternative to stuffing everything into the context window: amortize behavior and facts into weights instead of re-reading a prompt each turn. Whether it actually beats in-context learning and RAG is the open question the paper says it will test.

Browse previous days →