Astra goes GA, priced and flagged critical
OpenAI's GPT-6 Astra moved from limited preview to broad availability, arriving with public API pricing and the company's first-ever "critical" cybersecurity risk label. Around it the open-weight world kept shipping — compact long-context models from XHToken and document-retrieval models from Tencent — while hardware challengers and school regulators drew their own lines. A day of consolidation after last week's launch storm.
GPT-6 Astra reaches general availability at $10/$50 per million tokens
OpenAI began the broad rollout of GPT-6 Astra, priced at $10 per million input tokens and $50 per million output, available via the API and AWS now and to ChatGPT Plus, Pro, Business and Enterprise over the coming days. OpenAI designated Astra its first model rated a 'critical' cybersecurity risk, reporting a perfect 100% on ExploitBench, and gated the strongest cyber capabilities to select testing partners. Reported benchmarks include 97.6% on FrontierMath Tier 4 and 96.0% on GPQA Diamond, but a lower 57.2% on Humanity's Last Exam with tools. The system card also notes Astra's chain-of-thought monitorability decreased relative to GPT-5.6 Sol.
Why it matters: Concrete pricing plus API and AWS access mean developers can build on Astra today — but the critical-risk designation and reduced monitorability are the caveats to weigh before you do.
- OpenAI officially launches GPT-6 Astra: How to try it (Mashable)
- Nvidia's Huang says 'AGI has arrived' after OpenAI's GPT-6 Astra launch (Investing.com)
- Harvey + Legora on OpenAI's GPT-6 Astra (Artificial Lawyer)
XHToken's Spark-X2.5 pitches 1M-token context in 4B and 1.7B open models
A llama.cpp pull request surfaced Spark-X2.5-4B and Spark-X2.5-1.7B, compact open models from XHToken. Per the developer's own release notes, they use a hybrid attention design — one full-attention layer to three sliding-window layers — for native context up to 1M tokens and 200-plus languages, and integrate with the Codex, Claude Code and OpenClaw harnesses. The models were reportedly trained on Huawei Ascend clusters. No independent benchmarks are available yet.
Why it matters: Sub-5B models with a 1M-token window would be a cheap local option for long-context work — but the claims rest entirely on a single vendor's announcement.
Tencent releases EVIE visual-document retrieval models, claims ViDoRe V3 lead
Tencent published EVIE-8B and EVIE-4.5B on Hugging Face, open multi-vector embedding models for visual document retrieval. The model cards claim 66.75 nDCG@10 on ViDoRe V3 for the 8B and 66.02 for the 4.5B, with a training-free hierarchical clustering step compressing roughly 750 tokens per page down to 32 vectors to shrink the index. Tencent says the models were validated across 138 tasks spanning ViDoRe V1–V3 and JinaVDR; there is no third-party verification yet.
Why it matters: Retrieval over layouts, tables and charts is a stubborn RAG pain point, and open weights with a compact index make EVIE worth testing in document pipelines.
Preferred Networks lines up a 2028-2030 IPO to mass-produce MN-Core chips
Japan's Preferred Networks plans an IPO between 2028 and 2030 to fund volume production of its custom MN-Core inference processors, which the company claims run generative AI workloads up to 10x faster than conventional Nvidia GPUs — a figure it has not had independently verified. The chips use 3D-stacked memory to target bandwidth, the real inference bottleneck, and are fabbed by TSMC and Samsung. Backers include Toyota, SBI, Fanuc and NTT; the firm is valued around $2 billion with about 450 staff.
Why it matters: Another memory-bandwidth-first challenger to Nvidia on inference, though the headline 10x claim is the vendor's own and volume shipping is years out.
- Preferred Networks AI Chips IPO Targets Faster Generative AI Hardware (The Cryptonomist)
US federal government to pilot AI agents in job interviews
Per a CBS report, the US government will begin using AI virtual agents to run early-round interviews and screen applications for its two-year 'Tech Force' recruiting program, via the CodeSignal platform. Agents will handle phone, audio and text interviews, with hiring managers receiving transcribed recordings. The Office of Personnel Management has issued guidance urging agencies to use AI in hiring with human oversight, especially on crucial decisions.
Why it matters: A 1.9-million-employee public employer normalizing agent-led screening sets a template that other large employers — and candidates — will have to reckon with.
New York City bars student-facing generative AI through eighth grade
NYC Public Schools imposed a one-year moratorium on student-facing generative AI for grades 2-K through 8, affecting nearly 600,000 students, with companion-style chatbots banned across all grade levels. High schoolers get supervised, limited access plus two AI-literacy modules, and up to 50,000 can join approved pilots including Quill, Edia, Brisk Teaching, Playlab and Intel AI-Ready Schools. Teachers may use approved AI for lesson planning but are barred from using it to grade.
Why it matters: The largest US school district drawing a hard line on classroom AI creates a reference point that regulators and edtech vendors will track closely.
Phonely ships Alma, a voice-specific LLM undercutting GPT-4.1 on latency and price
Phonely launched Alma, an LLM trained on 10 million real phone conversations and built for voice agents, now available beyond its own platform. The company claims sub-185ms time-to-first-token versus roughly 500ms for GPT-4.1, and 55 cents per blended million tokens — which it frames as 84% cheaper than GPT-4.1. Alma works with any transcriber and text-to-speech provider and is designed for interruptions, background voices and transcription errors rather than clean turn-taking.
Why it matters: A narrow, domain-specific model beating a general frontier model on the two metrics that matter for telephony — latency and cost — is the kind of specialization voice-agent builders should watch.
Forensic audit of 8 abliterated Qwen 3.8 27B variants finds surgical edits win
A community project (abliterlitics.dev) benchmarked eight uncensored Qwen 3.8 27B variants over roughly 167 GPU hours using weight diffs, KL divergence, 13 benchmarks and HarmBench. The author reports the two smallest verified edits topped the refusal-removal rankings, while the most aggressive edit — 841 of 850 tensors touched — degraded capability and left 45% of adversarial responses looping past their token budget. The write-up also flags one variant shipping a 1,457-character jailbreak hidden inside its chat template. All figures are self-reported.
Why it matters: A rare adversarial audit of 'uncensored' model claims, and a concrete reminder to inspect chat templates, not just weights, before trusting a modified release.
Also worth a look
- Coding benchmarks that are quickly showcasing deep capability (Program-Bench, SRE-Bench, Code Migration) (r/LocalLLaMA)
- LayerStoRm open-source expert streaming: 1M context GLM-5.3-Flash on 96 GB VRAM via pinned experts in RAM (r/LocalLLaMA)
- Supporting independent journalism in Ukraine (OpenAI)
- The purpose of DNS is to spread scams (Simon Willison)
- I built an LLM benchmark harness that lets you browse and compare how models answered each question (r/LocalLLaMA)
- UH launches free 'AI for Hawaiʻi' course for everyone (University of Hawaii System)