Data & training
25 stories on this topic, newest first.
Allen AI open-sources Olmo-core 3, a trillion-parameter MoE training stack
Allen AI released Olmo-core 3, a redesigned open MoE training framework that keeps experts resident on GPUs via distributed data parallelism instead of repeatedly gathering weights under FSDP. In a preliminary 47B-parameter MoE test on eight NVIDIA B300s it processed 52,000 tokens/sec/GPU versus 19,400 for the old implementation, about 2.7x, and the stack has been benchmarked up to a 1.2-trillion-parameter model across 512 GPUs at 858 TFLOP/s/GPU. MXFP8 support added roughly 21 percent throughput over BF16 in a controlled run; the next Olmo will be MoE, and the stack is on GitHub with a technical report.
Why it matters: Open MoE training infrastructure, not just open weights, is what lets smaller labs train frontier-scale sparse models without reverse-engineering a proprietary stack.
NVIDIA open-sources Kumo Tabular, a tabular foundation model that tops four boards
NVIDIA released Kumo Tabular, an open foundation model for tabular classification and regression that predicts labels for new rows in a single forward pass via in-context learning, with no training, tuning or feature engineering. It was pretrained entirely on synthetic tables sampled from structural causal models, ships in three sizes (28M to 215M parameters) under the commercial-use OpenMDW-1.1 license, and NVIDIA says it ranks first on TabArena, BeyondArena, TALENT and ScoringBench while running about 17x faster than LimiX-2 on a single RTX 6000 Pro. Weights and a GPU-native library are on Hugging Face and GitHub.
Why it matters: Tabular prediction is the most common ML task in industry and has been gradient-boosted-tree territory for two decades. A drop-in, no-training foundation model with an open commercial license is a genuine shift in that workflow — if the leaderboard wins hold up on your own data.
Judge lets most of Reddit's scraping suit against Anthropic proceed
A San Francisco Superior Court judge allowed three of Reddit's five claims against Anthropic to move forward — breach of contract, interference with contract, and California unfair competition — while dismissing unjust enrichment and trespass to chattels with leave to refile by Oct 16. The judge rejected Anthropic's argument that Reddit's terms were an unenforceable "browsewrap," citing allegations that Anthropic kept scraping the site over 100,000 times after Reddit's CEO publicly objected.
Why it matters: A contract-law route to holding model trainers liable for scraping, distinct from the copyright fights — and a signal that click-free terms of service may still bind crawlers.
Basecamp Research raises $140M to turn wild DNA into training data
London's Basecamp Research raised $140 million, with S32 leading and Nvidia, Anthropic's Anthology Fund and the NATO Innovation Fund taking part, to expand its EDEN biological models trained on genetic material collected from rainforests, hot springs and the deep sea. CTO Philip Lorenz says the roughly 15-trillion-token dataset — where a token is a single DNA base — is meant to grow about 100x toward the Trillion Gene Atlas being built with Anthropic, Nvidia, PacBio and Ultima Genomics. Basecamp reports an EDEN-designed antibiotic, EDEN-7, matched a last-resort drug against resistant bacteria in mice, and that 97% of tested antimicrobial peptides showed lab activity. It won't open-source the models, citing biosecurity.
Why it matters: It's a concrete answer to the 'where does more data come from' question — not the web, but nature — and a reminder that scaling laws are being stress-tested well outside language.
Shopify distills a production agent past frontier quality at 4% of the serving cost
In a PyTorch case study, Shopify details a continual-learning loop that mines anonymized production failures, has a panel of frontier models 'heal' them into training trajectories, and folds them back into a smaller model via supervised fine-tuning and GRPO. Its GraphQL agent, serving up to 2,000 requests per minute, is claimed to beat the frontier baseline while cutting serving cost roughly 96% — from an estimated $27M to about $1M per year. Gist-token compression shrank the static system prompt from ~6,000 to ~1,500 tokens, dropping time-to-first-token ~19% and end-to-end latency ~38% under load. The judge is calibrated against human annotators with DSPy optimizers, and inference runs on vLLM.
Why it matters: This is the most concrete published recipe yet for the 'frontier to launch, distill to scale' pattern — the same economics driving the labs' own price cuts, but done in-house against your own traffic.
Unsealed NYT filings quote Microsoft calling AI scraping 'the largest theft of labor in human history'
A newly unsealed summary-judgment brief in the New York Times' three-year-old suit against OpenAI and Microsoft surfaces internal documents the companies had kept confidential. In a January 2023 memo, Microsoft applied-science director Brent Hecht called the training practice 'an astonishing theft of unprecedented proportions' and 'the largest theft of labor in human history.' The filing cites specifics: OpenAI mid-training datasets allegedly holding 91,692 copies of NYT, Daily News and CIR works; a Common Crawl-derived set with over 2 million nytimes.com documents; and Copilot cutting click-through to the NYT domain by as much as 93% versus Bing search. Many quotes come from the plaintiffs' own brief, stripped of original context; the underlying exhibits remain sealed, and OpenAI and Microsoft did not comment.
Why it matters: The admissions cut directly at the fair-use defense the industry is leaning on, particularly the market-harm prong, and the Trump administration filed in OpenAI's defense earlier this month. If they survive context, they reshape the leverage in every training-data suit.
AfterQuery becomes YC's fastest unicorn at a $3.2B valuation
AI training-data startup AfterQuery reportedly raised a round valuing it at $3.2 billion, five months after announcing a $30 million Series A at a $300 million valuation — which YC partner Gustaf Alströmer calls the accelerator's fastest ever launch-to-unicorn. In April the company reported a $100 million annualized revenue run rate and named Nvidia, Legora, and Motif Technologies as customers. Rather than optimizing answer accuracy, AfterQuery trains models and agents to replicate how professionals like doctors and lawyers complete tasks. Forbes first reported the round.
Why it matters: The training-data layer — Mercor, Scale, now AfterQuery — keeps commanding frontier-scale valuations, a signal that expert task data, not just more compute, is the current bottleneck labs pay up for.
LAION releases 10-million-hour open video dataset
LAION published the Big Video Dataset (BVD), drawn from 1.3 billion video URLs in CommonCrawl. It downloaded 80 million videos totaling 10 million hours, extracting 55 million clips with auto-generated video and audio descriptions plus 300 million still images. LAION says models trained on BVD outperform comparable InternVid-trained models by up to 2.1 percentage points on video-to-text benchmarks. The dataset is research-only, with LAION leaning on a 2024 Hamburg court ruling permitting collection of copyrighted content for non-commercial research.
Why it matters: One of the largest openly available video corpora lowers the barrier to training multimodal and world models, but the research-only framing and copyright basis leave commercial use in a legal gray zone.
Lawsuit alleges xAI trained Grok on child sexual abuse material
A complaint filed by a plaintiff known as Jane Doe alleges xAI trained Grok on child sexual abuse material (CSAM), after the Canadian Centre for Child Protection notified her that AI-generated CSAM depicting her was identified on xAI. Her images had been hashed decades ago by NCMEC and the CCCP. The suit cites forum messages among offenders discussing the creation of AI-generated CSAM of known legacy victims.
Why it matters: Training-data provenance is moving from abstract copyright disputes to criminal-grade liability, putting xAI's data pipeline and content filtering squarely before a court.
- Elon Musk's xAI used child porn to train Grok models, lawsuit says (Ars Technica AI)
The data-efficiency gap: kids learn language on a rounding error of an LLM's tokens
MIT Technology Review surveys the BabyLM effort and the 'data efficiency gap': a preteen hears roughly 100 million words, versus the 15 trillion tokens Llama 3.1 pretrained on. The 2024 champion GPT-BERT, trained on about 100 million words, still beat Llama 2 70B on one BabyLM benchmark, while popular ideas like curriculum learning underperformed and multimodal training on baby-headcam video remains weak. With easily available web text possibly running dry by the 2030s, efficiency is becoming the constraint.
Why it matters: If pretraining data is finite, learning more from less is the next real frontier — and the most plausible way universities and minority-language communities stay in the game.
- Kids outlearn AI—and we still don't know why (MIT Technology Review)
Chinese 'transfer stations' resell Claude tokens at 10% of list price
An Oxford China Policy Lab analysis details a modular supply chain of API proxies — 'transfer stations' — that route Chinese developers' requests through overseas servers, defeating Anthropic's geoblocking, KYC and biometric checks. Operators farm free credits, split Max plans, and quietly 'dilute' requests by swapping Opus for Sonnet or Chinese models; researchers found one fake 'Gemini-2.5' endpoint scoring 37% on a medical benchmark versus the official 84%. The likely real prize is the logs — prompts and tool calls harvested for distillation, with Claude Opus 4.6 reasoning traces already circulating on Hugging Face.
Why it matters: The same infrastructure that beats export controls also blinds abuse-monitoring systems like Clio — and if you buy tokens through a proxy, your prompts may become someone's training set.
AirTag traces Amazon's bulk rare-book buys to a book-shredding AI scanning lab
404 Media convinced a bookseller to plant an Apple AirTag in a ~1,000-book bulk order; it ended up at Amazon's VGT3 team inside the LAS8 facility in Las Vegas, whose door logo is a T. rex devouring a book. Workers there cut spines off books to speed destructive scanning, and Amazon uses the pages to train its Nova models. Amazon's statement said only that it 'purchases books through commercial channels'; the practice mirrors Anthropic's court-revealed 'Project Panama.'
Why it matters: Pre-2022 printed text is now a scarce, contamination-free training asset worth destroying originals for — the data land grab has moved from scraping the web to physically shredding the archive.
- We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility (Simon Willison)
- Amazon, which started off selling books, is destroying rare texts to train AI (TechCrunch AI)
- Hidden Airtag reveals Amazon is trashing rare books to train AI (Ars Technica AI)
- AirTag reveals how Amazon destroys rare books for AI training (The Decoder)
LittleLearner: a model that never learns past fifth grade
Researchers trained 0.6B-5B models from scratch on LittleCurriculum, an 88B-token corpus filtered to the US K-5 curriculum, alongside matched unfiltered controls. Across scaling, SFT+GRPO post-training, and in-context learning, every intervention amplified in-scope ability but none meaningfully improved out-of-scope performance — the pretraining filter set a hard capability ceiling. A 5B chat model is live in-browser.
Why it matters: It's a clean experimental handle on the 'learned vs merely elicited' question: if capabilities can't be coaxed past what the pretraining data contained, that bounds what RL and prompting can realistically unlock.
World Labs turns one robot demo into thousands of simulated variations
Fei-Fei Li's World Labs unveiled a Real-to-Sim-to-Real engine that rebuilds a real robot task as a physically faithful interactive world, then generates thousands of variations — lighting, object count, friction, camera angle — to train control policies entirely in simulation. Policies trained only in sim ran for an hour each on four robot platforms without human intervention, and model rankings in sim matched reality across GR00T N1.6 and π₀.₅ checkpoints, evaluated with 2,000 simulated versus 100 real runs each.
Why it matters: If sim evaluation reliably predicts which policy wins on hardware, robotics gets the fast iteration loop LLMs already enjoy — cutting the expensive real-world testing that has held the field back.
Twitch opts every streamer into Amazon AI training by default
Twitch added a setting letting users opt out of having their streams, VODs, clips, chats, and channel text used to train Amazon's generative AI models, defaulting everyone to opted-in. Asked why it isn't opt-in, CPO Mike Minton said on stream: "if it was opt-in, nobody would opt in." A user-forum request to reverse the default has topped 13,000 upvotes. One mitigation: Twitch auto-deletes VODs at 60 days, capping what Amazon can pull to a streamer's most recent window.
Why it matters: It's a rare on-the-record admission of the opt-out playbook platforms use to convert user content into training data, and a reminder to check the default consent settings on anything you host.
- Amazon will train on Twitch streamers' content by default, unless they opt out (TechCrunch AI)
- "If it was opt-in, nobody would opt in": Twitch auto-enrolls all streamers in Amazon LLM training (tubefilter.com)
- 'If it was opt in, nobody would opt in': Twitch is using streamers' content to train generative AI by default (Video Games Chronicle)
- Twitch content has trained Amazon AI for years, but users can opt out now (Ars Technica AI)
FineBooks benchmarks OCR models to salvage public-domain training data
Hugging Face and EleutherAI's FineBooks project tested 14 open-weight OCR models on 2,165 historical book pages with expert ground truth, publishing a leaderboard scored by character error rate. Old OCR is a real training tax: the Talkie project found models learn at only 30% efficiency on OCR text versus clean human transcriptions. The best models now clear 97% character accuracy at under $2 per 1,000 pages, and size doesn't track quality, the 3B dots.ocr tops the 9B Qwen3.5, and a 0.9B model takes second. The team plans to reprocess ~200,000 public-domain Biodiversity Heritage Library documents and release the cleaned text.
Why it matters: Reprocessing the 300K-book Common Pile with modern OCR is one of the cheapest ways to improve openly licensed pretraining corpora. The catch: these models silently modernize archaic characters, so they're good enough for training but not for scholarship.
Notion open-sources Zerank 2, giving local RAG a SOTA reranker
A practitioner benchmark for a 15-language translation-memory retrieval task found F2LLM V2 4B embeddings paired with Zerank 2 4B reranking (0.919 MRR, 98.4% recall@20) beating Qwen 3, BGE-M3, and even Voyage 4 Large plus Voyage Rerank 2.5 over API. Both models are fully open: F2LLM ships open weights, data, and code, and Zerank 2 was released under a permissive license after Notion acquired ZeroEntropy 16 days ago.
Why it matters: A fully open, self-hostable embedding-plus-reranker stack that edges out paid API rerankers is a concrete upgrade path for anyone running RAG without shipping queries to a vendor.
- Best Embedding + Reranking Model (r/LocalLLaMA)
ByteDance pre-trains a 10-trillion-parameter model to chase Mythos
Per the Financial Times, ByteDance is early in pre-training a model with as many as 10 trillion parameters — three times Moonshot's Kimi K3 and in the range of estimates for Anthropic's ~8T Mythos 5. Sources say ByteDance has avoided distillation from rival model outputs for over a year, and founder Zhang Yiming has told the 2,000-person Seed team to aim for world-leading capability. xAI is reportedly training 6T and 10T Grok variants on its Colossus 2 cluster.
Why it matters: The parameter gap between Chinese labs and the US frontier is closing fast, and raw scale is back in fashion at the very moment everyone else is preaching the efficiency frontier.
- ByteDance trains massive AI model in bid to rival Anthropic (Ars Technica)
- China's Largest AI Model Is Being Developed at Bytedance (The Decoder)
Audit finds ~12% of GPQA, MMLU-Pro and MMMU-Pro questions broken
A community audit of GPQA (Diamond and Extended), MMLU-Pro and MMMU-Pro found roughly 12% of questions verifiably broken — malformed, with wrong answer keys, or with more than one defensible answer. After cleaning, top models jump from the ~92–93% ceiling on GPQA-Diamond to around 98%, implying the plateau was the benchmark, not the models. The author released -Clean versions of all four benchmarks, a flagged-candidate ledger, lm-eval-harness tasks and Hugging Face datasets.
Why it matters: If a tenth of your eval is wrong, 'near-saturation' scores are noise — and since the corrected sets and the ledger are public, there's no excuse to keep quoting the dirty numbers.
Hugging Face ships The Stack v3, a 114TB open code corpus
Hugging Face released The Stack v3, its largest open code dataset yet. It comes in two forms: stack-v3-train, a near-deduplicated, quality-filtered, PII-redacted set with contents inline for immediate load_dataset use; and stack-v3-full, the entire 114TB corpus as an HF storage bucket with every duplicate kept and cluster IDs, for teams that want to roll their own dedup, filters and mixes.
Why it matters: An openly licensed code pretraining corpus at this scale is rare fuel for anyone training or fine-tuning coding models outside the big labs.
Xaira bets causal CRISPR data, not scale, unlocks the virtual cell
On Latent Space, Xaira's Ci Chu and Bo Wang argue that RNA-expression 'virtual cell' models trained on correlational data like CELLxGENE plateau — a 3.1B model falls off the scaling curve because the data is information-limited, not compute-limited. Their fix is X-Atlas, built from millions of parallel CRISPR perturbation experiments that knock genes down one at a time to capture causal upstream/downstream effects, roughly 30x more information, which restores parameter and compute scaling for their X-Cell model. They also abandoned autoregression for diffusion.
Why it matters: It's a clean illustration of the data-vs-scale ceiling: when test loss flatlines, more parameters won't help, and building the right causal dataset is the actual lever — a lesson that generalizes well beyond biology.
- Causal Models Need Causal Data — Xaira's X-Cell model for Drug Discovery (Latent Space (swyx))
Judge signs off on Anthropic's $1.5B book-piracy settlement
US District Judge Araceli Martinez-Olguin granted final approval to Anthropic's $1.5 billion class-action settlement, paying roughly $3,000 per work across about 500,000 titles it downloaded from pirate libraries like Library Genesis to train Claude. The late Judge Alsup's underlying ruling stands: training on copyrighted text is fair use, but obtaining it via piracy is not, and Anthropic must now destroy the pirated copies. Because Anthropic settled rather than appealed, none of this becomes binding precedent, and parallel suits against Google, Meta, OpenAI and Midjourney roll on.
Why it matters: Fair-use-for-training survives as the industry's working assumption, but provenance is now a nine-to-ten-figure liability: where you sourced the data matters as much as what you did with it.
- Anthropic's landmark $1.5B copyright settlement is approved (TechCrunch AI)
- Judge Approves Anthropic's Record-Breaking $1.5 Billion Settlement For AI Copyright Lawsuit (Engadget)
- Anthropic settles with authors and publishers for $1.5B in landmark copyright case (SiliconANGLE)
- US judge approves Anthropic's $1.5 billion settlement of copyright lawsuit (Reuters)
Google's SensorFM: one foundation model for wearable sensor data
Google Research unveiled SensorFM, a foundation model pretrained self-supervised on over a trillion minutes of unlabeled Fitbit and Pixel Watch data from five million people across 100+ countries. It processes 34 features from five sensor types (PPG, acceleration, skin conductance and temperature, altitude) and beat supervised baselines with hand-crafted features on 34 of 35 downstream health tasks. Performance scaled cleanly with model and data size, from ~100K to 100M parameters. It remains research-only, aggregated to minute-level data, and tested only on Google's own devices.
Why it matters: It's the wearables version of the 'one big pretrained model replaces many task-specific ones' pattern, and a signal for where personal-health agents get their context. Note the caveats: no raw signals, self-reported labels, and no shipping plans.
sqlite-utils 4.0rc1 adds migrations and nested transactions
Simon Willison released the first release candidate for sqlite-utils v4, folding the proven sqlite-migrate package in directly as a built-in migrations system driven by decorated Python functions and a new migrate CLI command. It also adds db.atomic() for nested transactions backed by SQLite savepoints, borrowing Django/Peewee terminology. The major bump carries breaking changes: type detection now defaults on for CSV/TSV import, REAL replaces FLOAT, schemas use double-quotes, and db.table() no longer returns views.
Why it matters: A widely used building block for LLM data pipelines gets first-class migrations and transactions — worth testing the breaking changes before the stable release lands.
- sqlite-utils 4.0rc1 adds migrations and nested transactions (Simon Willison)
- sqlite-utils 4.0rc1 (Simon Willison)
Fine-tuning Qwen 3 0.6B turns a tiny model into a 92%-accurate classifier
A developer building a household RAG chatbot fine-tuned Qwen 3 0.6B with Unsloth and QLoRA to categorize incoming questions and narrow the vector search space. Prompting the base model alone scored just 10% on a 131-test battery; fine-tuning lifted it to 79%. Mapping categories to two-character opaque IDs with no semantic overlap — instead of human-readable labels — pushed accuracy to ~92% by eliminating fragment and confusion errors.
Why it matters: A concrete reminder that a 600M-parameter local model can handle narrow classification reliably after fine-tuning, and that output-format design often beats prompt-tweaking.