<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><title>gonioAI — Data &amp; training</title><link>https://gonioai.pages.dev/topics/data/</link><description>Data &amp; training stories from gonioAI.</description><language>en</language><lastBuildDate>Tue, 11 Aug 2026 10:45:13 +0000</lastBuildDate><item><title>FineBooks benchmarks OCR models to salvage public-domain training data</title><link>https://the-decoder.com/old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at-scale</link><guid isPermaLink="false">2026-08-11:data:https://the-decoder.com/old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at-scale</guid><pubDate>Tue, 11 Aug 2026 07:00:00 +0000</pubDate><description>Hugging Face and EleutherAI's FineBooks project tested 14 open-weight OCR models on 2,165 historical book pages with expert ground truth, publishing a leaderboard scored by character error rate. Old OCR is a real training tax: the Talkie project found models learn at only 30% efficiency on OCR text versus clean human transcriptions. The best models now clear 97% character accuracy at under $2 per 1,000 pages, and size doesn't track quality, the 3B dots.ocr tops the 9B Qwen3.5, and a 0.9B model takes second. The team plans to reprocess ~200,000 public-domain Biodiversity Heritage Library documents and release the cleaned text.

Why it matters: Reprocessing the 300K-book Common Pile with modern OCR is one of the cheapest ways to improve openly licensed pretraining corpora. The catch: these models silently modernize archaic characters, so they're good enough for training but not for scholarship.</description></item><item><title>Notion open-sources Zerank 2, giving local RAG a SOTA reranker</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1vjk57h/best_embedding_reranking_model</link><guid isPermaLink="false">2026-08-09:data:https://www.reddit.com/r/LocalLLaMA/comments/1vjk57h/best_embedding_reranking_model</guid><pubDate>Sun, 09 Aug 2026 07:00:00 +0000</pubDate><description>A practitioner benchmark for a 15-language translation-memory retrieval task found F2LLM V2 4B embeddings paired with Zerank 2 4B reranking (0.919 MRR, 98.4% recall@20) beating Qwen 3, BGE-M3, and even Voyage 4 Large plus Voyage Rerank 2.5 over API. Both models are fully open: F2LLM ships open weights, data, and code, and Zerank 2 was released under a permissive license after Notion acquired ZeroEntropy 16 days ago.

Why it matters: A fully open, self-hostable embedding-plus-reranker stack that edges out paid API rerankers is a concrete upgrade path for anyone running RAG without shipping queries to a vendor.</description></item><item><title>ByteDance pre-trains a 10-trillion-parameter model to chase Mythos</title><link>https://arstechnica.com/ai/2026/08/bytedance-trains-massive-ai-model-in-bid-to-rival-anthropic</link><guid isPermaLink="false">2026-08-08:data:https://arstechnica.com/ai/2026/08/bytedance-trains-massive-ai-model-in-bid-to-rival-anthropic</guid><pubDate>Sat, 08 Aug 2026 07:00:00 +0000</pubDate><description>Per the Financial Times, ByteDance is early in pre-training a model with as many as 10 trillion parameters — three times Moonshot's Kimi K3 and in the range of estimates for Anthropic's ~8T Mythos 5. Sources say ByteDance has avoided distillation from rival model outputs for over a year, and founder Zhang Yiming has told the 2,000-person Seed team to aim for world-leading capability. xAI is reportedly training 6T and 10T Grok variants on its Colossus 2 cluster.

Why it matters: The parameter gap between Chinese labs and the US frontier is closing fast, and raw scale is back in fashion at the very moment everyone else is preaching the efficiency frontier.</description></item><item><title>Audit finds ~12% of GPQA, MMLU-Pro and MMMU-Pro questions broken</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1v99f6m/paper_gpqa_mmlupro_and_mmmupro_were_audited_for</link><guid isPermaLink="false">2026-07-29:data:https://www.reddit.com/r/LocalLLaMA/comments/1v99f6m/paper_gpqa_mmlupro_and_mmmupro_were_audited_for</guid><pubDate>Wed, 29 Jul 2026 07:00:00 +0000</pubDate><description>A community audit of GPQA (Diamond and Extended), MMLU-Pro and MMMU-Pro found roughly 12% of questions verifiably broken — malformed, with wrong answer keys, or with more than one defensible answer. After cleaning, top models jump from the ~92–93% ceiling on GPQA-Diamond to around 98%, implying the plateau was the benchmark, not the models. The author released -Clean versions of all four benchmarks, a flagged-candidate ledger, lm-eval-harness tasks and Hugging Face datasets.

Why it matters: If a tenth of your eval is wrong, 'near-saturation' scores are noise — and since the corrected sets and the ledger are public, there's no excuse to keep quoting the dirty numbers.</description></item><item><title>Hugging Face ships The Stack v3, a 114TB open code corpus</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1v59aek/hugging_face_releases_the_stack_v3_largest_open</link><guid isPermaLink="false">2026-07-25:data:https://www.reddit.com/r/LocalLLaMA/comments/1v59aek/hugging_face_releases_the_stack_v3_largest_open</guid><pubDate>Sat, 25 Jul 2026 07:00:00 +0000</pubDate><description>Hugging Face released The Stack v3, its largest open code dataset yet. It comes in two forms: stack-v3-train, a near-deduplicated, quality-filtered, PII-redacted set with contents inline for immediate load_dataset use; and stack-v3-full, the entire 114TB corpus as an HF storage bucket with every duplicate kept and cluster IDs, for teams that want to roll their own dedup, filters and mixes.

Why it matters: An openly licensed code pretraining corpus at this scale is rare fuel for anyone training or fine-tuning coding models outside the big labs.</description></item><item><title>Xaira bets causal CRISPR data, not scale, unlocks the virtual cell</title><link>https://www.latent.space/p/xaira</link><guid isPermaLink="false">2026-07-22:data:https://www.latent.space/p/xaira</guid><pubDate>Wed, 22 Jul 2026 07:00:00 +0000</pubDate><description>On Latent Space, Xaira's Ci Chu and Bo Wang argue that RNA-expression 'virtual cell' models trained on correlational data like CELLxGENE plateau — a 3.1B model falls off the scaling curve because the data is information-limited, not compute-limited. Their fix is X-Atlas, built from millions of parallel CRISPR perturbation experiments that knock genes down one at a time to capture causal upstream/downstream effects, roughly 30x more information, which restores parameter and compute scaling for their X-Cell model. They also abandoned autoregression for diffusion.

Why it matters: It's a clean illustration of the data-vs-scale ceiling: when test loss flatlines, more parameters won't help, and building the right causal dataset is the actual lever — a lesson that generalizes well beyond biology.</description></item><item><title>Judge signs off on Anthropic's $1.5B book-piracy settlement</title><link>https://techcrunch.com/2026/07/20/anthropics-landmark-1-5b-copyright-settlement-is-approved</link><guid isPermaLink="false">2026-07-21:data:https://techcrunch.com/2026/07/20/anthropics-landmark-1-5b-copyright-settlement-is-approved</guid><pubDate>Tue, 21 Jul 2026 07:00:00 +0000</pubDate><description>US District Judge Araceli Martinez-Olguin granted final approval to Anthropic's $1.5 billion class-action settlement, paying roughly $3,000 per work across about 500,000 titles it downloaded from pirate libraries like Library Genesis to train Claude. The late Judge Alsup's underlying ruling stands: training on copyrighted text is fair use, but obtaining it via piracy is not, and Anthropic must now destroy the pirated copies. Because Anthropic settled rather than appealed, none of this becomes binding precedent, and parallel suits against Google, Meta, OpenAI and Midjourney roll on.

Why it matters: Fair-use-for-training survives as the industry's working assumption, but provenance is now a nine-to-ten-figure liability: where you sourced the data matters as much as what you did with it.</description></item><item><title>Google's SensorFM: one foundation model for wearable sensor data</title><link>https://the-decoder.com/sensorfm</link><guid isPermaLink="false">2026-07-13:data:https://the-decoder.com/sensorfm</guid><pubDate>Mon, 13 Jul 2026 07:00:00 +0000</pubDate><description>Google Research unveiled SensorFM, a foundation model pretrained self-supervised on over a trillion minutes of unlabeled Fitbit and Pixel Watch data from five million people across 100+ countries. It processes 34 features from five sensor types (PPG, acceleration, skin conductance and temperature, altitude) and beat supervised baselines with hand-crafted features on 34 of 35 downstream health tasks. Performance scaled cleanly with model and data size, from ~100K to 100M parameters. It remains research-only, aggregated to minute-level data, and tested only on Google's own devices.

Why it matters: It's the wearables version of the 'one big pretrained model replaces many task-specific ones' pattern, and a signal for where personal-health agents get their context. Note the caveats: no raw signals, self-reported labels, and no shipping plans.</description></item><item><title>sqlite-utils 4.0rc1 adds migrations and nested transactions</title><link>https://simonwillison.net/2026/Jun/21/sqlite-utils-40rc1</link><guid isPermaLink="false">2026-06-22:data:https://simonwillison.net/2026/Jun/21/sqlite-utils-40rc1</guid><pubDate>Mon, 22 Jun 2026 07:00:00 +0000</pubDate><description>Simon Willison released the first release candidate for sqlite-utils v4, folding the proven sqlite-migrate package in directly as a built-in migrations system driven by decorated Python functions and a new migrate CLI command. It also adds db.atomic() for nested transactions backed by SQLite savepoints, borrowing Django/Peewee terminology. The major bump carries breaking changes: type detection now defaults on for CSV/TSV import, REAL replaces FLOAT, schemas use double-quotes, and db.table() no longer returns views.

Why it matters: A widely used building block for LLM data pipelines gets first-class migrations and transactions — worth testing the breaking changes before the stable release lands.</description></item><item><title>Fine-tuning Qwen 3 0.6B turns a tiny model into a 92%-accurate classifier</title><link>https://www.teachmecoolstuff.com/viewarticle/fine-tuning-a-local-llm-to-categorize-questions</link><guid isPermaLink="false">2026-06-22:data:https://www.teachmecoolstuff.com/viewarticle/fine-tuning-a-local-llm-to-categorize-questions</guid><pubDate>Mon, 22 Jun 2026 07:00:00 +0000</pubDate><description>A developer building a household RAG chatbot fine-tuned Qwen 3 0.6B with Unsloth and QLoRA to categorize incoming questions and narrow the vector search space. Prompting the base model alone scored just 10% on a 131-test battery; fine-tuning lifted it to 79%. Mapping categories to two-character opaque IDs with no semantic overlap — instead of human-readable labels — pushed accuracy to ~92% by eliminating fragment and confusion errors.

Why it matters: A concrete reminder that a 600M-parameter local model can handle narrow classification reliably after fine-tuning, and that output-format design often beats prompt-tweaking.</description></item></channel></rss>
