<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><title>gonioAI — Multimodal</title><link>https://gonioai.pages.dev/topics/multimodal/</link><description>Multimodal stories from gonioAI.</description><language>en</language><lastBuildDate>Tue, 11 Aug 2026 10:45:13 +0000</lastBuildDate><item><title>A $2,000 connector gives frozen DeepSeek V4 Flash basic vision</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1vl6ior/i_gave_deepseek_v4_flash_basic_vision_by_training</link><guid isPermaLink="false">2026-08-11:multimodal:https://www.reddit.com/r/LocalLLaMA/comments/1vl6ior/i_gave_deepseek_v4_flash_basic_vision_by_training</guid><pubDate>Tue, 11 Aug 2026 07:00:00 +0000</pubDate><description>A developer bolted vision onto text-only DeepSeek V4 Flash (284B total / 13B active) without touching the language model, freezing both it and a 417M MoonViT encoder and training only a 40.1M-parameter connector on 100K image-text examples (39,619 unique images). One epoch on 5x H200s, ~$2,000 end to end, produced a working NVFP4 model that reads storefront signs and grounds UI controls, though it still misses small text and hallucinates details. The recipe follows Baseten's frozen-MoE GLM-5.2 Vision work; the author estimates a production-grade 1M-example run at $15-20K and released weights for both the DeepSeek and a smaller Laguna XS 2.1 variant.

Why it matters: It's a cheap, reproducible template for retrofitting perception onto strong open text models instead of waiting for native VLMs, handy for anyone building browser or desktop agents that need to see screenshots. The bottleneck is now data scale, not the method.</description></item><item><title>MiniMax open-weights H3, a video model that generates its own audio</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1vkdatu/minimax_h3_a_new_openweight_video_model_live_in</link><guid isPermaLink="false">2026-08-10:multimodal:https://www.reddit.com/r/LocalLLaMA/comments/1vkdatu/minimax_h3_a_new_openweight_video_model_live_in</guid><pubDate>Mon, 10 Aug 2026 07:00:00 +0000</pubDate><description>MiniMax released H3, an open-weight multimodal video model now runnable in ComfyUI for text/image/video-to-video, first- and last-frame generation, and reference-driven creation. Unlike pipelines that dub audio afterward, H3 jointly generates visuals and synchronized stereo audio—dialogue, sound effects, ambience, and music—in one pass. Open checkpoints handle clips up to 15 seconds at 768p; MiniMax's hosted version goes up to 2K.

Why it matters: Joint audio-video generation in open weights is still rare. Local creators get a single-model pipeline instead of stitching a separate video model to a separate audio one.</description></item><item><title>xAI ships Imagine Image 2.0, lands #2 behind GPT-Image-2</title><link>https://the-decoder.com/xais-imagine-image-2-0-lands-just-behind-openais-gpt-image-2-in-arena-benchmarks</link><guid isPermaLink="false">2026-08-08:multimodal:https://the-decoder.com/xais-imagine-image-2-0-lands-just-behind-openais-gpt-image-2-in-arena-benchmarks</guid><pubDate>Sat, 08 Aug 2026 07:00:00 +0000</pubDate><description>xAI launched Imagine Image 2.0 as a 'Quality Mode' in Grok's web and mobile apps, adding a Magic Wand for localized edits, region segmentation, background removal, multi-reference editing (up to five inputs), and smart resize with generative fill. Its faster 'low' variant sits second on both Arena boards as of Aug 7 — 1,439 Elo in Image Edit and 1,320 in Text-to-Image — behind OpenAI's GPT-Image-2 (1,463 / 1,380) and ahead of Reve, Meta Muse-Image, Qwen-Image-3.0-Pro, Gemini and SeedDream. API access is 'coming soon.'

Why it matters: The image-model leaderboard is now a genuine multi-way scrum; GPT-Image-2 still sets the bar, but no longer sits alone at the top.</description></item><item><title>NVIDIA ships Cosmos 3, an open world-model family for physical AI</title><link>https://blogs.nvidia.com/blog/open-world-models-physical-ai</link><guid isPermaLink="false">2026-08-07:multimodal:https://blogs.nvidia.com/blog/open-world-models-physical-ai</guid><pubDate>Fri, 07 Aug 2026 07:00:00 +0000</pubDate><description>NVIDIA released Cosmos 3, a mixture-of-transformers 'omni' family under the OpenMDW 1.1 license that combines vision reasoning, world generation, and action prediction in one stack. It comes in three sizes: Super (64B), Nano (16B), and Edge (4B) for on-device robot policy on Jetson and RTX GPUs. NVIDIA claims top open-weights rankings on Artificial Analysis for text-to-image and image-to-video, plus No. 1 on RoboLab for robot policy.

Why it matters: World models that generate physically grounded synthetic data and simulate future states are the emerging substrate for robotics and AV teams, and open weights plus an Edge tier make specialization on your own hardware realistic.</description></item><item><title>Scenema Audio brings expressive voice cloning to ComfyUI on 8GB VRAM</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1vgfmee/scenema_audio_comes_to_comfyui_runs_on_8gb_vram</link><guid isPermaLink="false">2026-08-06:multimodal:https://www.reddit.com/r/LocalLLaMA/comments/1vgfmee/scenema_audio_comes_to_comfyui_runs_on_8gb_vram</guid><pubDate>Thu, 06 Aug 2026 07:00:00 +0000</pubDate><description>The text-to-speech model behind scenema.ai landed as a native ComfyUI custom node, quantized to run on 8GB VRAM (tested on RTX 3070 and 4090) at up to 2x realtime. It offers zero-shot voice cloning and inline stage-direction cues like [voice cracks] performed at the exact spot, replacing the original XML prompt format with bracket tags. Node code is MIT; the transformer weights derive from the LTX-2 Community License and use a gated Gemma 3 12B text encoder, with a one-time ~30GB weight download.

Why it matters: Diffusion-based expressive TTS with voice cloning is now self-hostable on a mid-range consumer GPU — a practical local alternative to cloud voice APIs, caveats about seed-dependent gibberish aside.</description></item><item><title>Mistral's Shieldstral makes content moderation a prompt, not a retrain</title><link>https://mistral.ai/news/shieldstral</link><guid isPermaLink="false">2026-08-05:multimodal:https://mistral.ai/news/shieldstral</guid><pubDate>Wed, 05 Aug 2026 07:00:00 +0000</pubDate><description>Mistral released Shieldstral, a 3B open-weights (Apache 2.0) multimodal safety classifier that frames moderation as policy-adaptive yes/no question answering: you supply a plain-language policy at inference time and get a calibrated safety score from a single forward pass. It handles text, images, and prompt-response pairs, runs on a single 16GB GPU, and Mistral claims it matches open guard models up to 7x larger on text safety while setting a new bar on multimodal moderation. vLLM shipped day-zero serving with one-forward-pass scoring, 12 languages, and 32k context.

Why it matters: Guardrail models that bake a fixed harm taxonomy into their weights force a retrain per deployment; a policy-in-the-prompt classifier that runs on one 16GB card is a far cheaper way to re-target moderation per product.</description></item><item><title>MiniMax H3 open weights land on Hugging Face</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1ve1mvh/minimaxh3_now_on_huggingface</link><guid isPermaLink="false">2026-08-03:multimodal:https://www.reddit.com/r/LocalLLaMA/comments/1ve1mvh/minimaxh3_now_on_huggingface</guid><pubDate>Mon, 03 Aug 2026 07:00:00 +0000</pubDate><description>MiniMax released open weights for H3, an omni-modal system that understands text, images, video and audio and generates video with native stereo audio at up to 2K resolution and 15-second durations. Early community comparisons pit its output against Seedance 2.5. The model was teased earlier in the week; the weights are now actually downloadable.

Why it matters: An open-weight video-plus-audio generator is a rare thing, and it drops the barrier for local video pipelines that previously meant a closed API subscription.</description></item><item><title>Google pulls Google Earth's AI image feature two days after launch</title><link>https://the-decoder.com/google-handed-users-the-easiest-possible-tool-for-fake-satellite-imagery-then-pulled-it-after-two-days</link><guid isPermaLink="false">2026-08-01:multimodal:https://the-decoder.com/google-handed-users-the-easiest-possible-tool-for-fake-satellite-imagery-then-pulled-it-after-two-days</guid><pubDate>Sat, 01 Aug 2026 07:00:00 +0000</pubDate><description>Google rolled out and then quickly retracted a Nano Banana 2 integration in Google Earth that let anyone generate custom scenes superimposed on real satellite, aerial and 3D imagery. Users immediately demonstrated fabricated refugee columns at the Mexican border and bombed-out hospitals, prompting Google to roll back the feature pending stronger guardrails. The company says generated images were labeled AI and not visible to other Earth users.

Why it matters: Google marketed a tool that made convincing geospatial disinformation trivially easy on a platform journalists treat as ground truth, a reminder that provenance labels are weak defense once a screenshot leaves the app.</description></item><item><title>MiniMax H3 undercuts video generators and promises open weights</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1vbdsmz/minimaxh3_video_model_released_open_weights</link><guid isPermaLink="false">2026-07-31:multimodal:https://www.reddit.com/r/LocalLLaMA/comments/1vbdsmz/minimaxh3_video_model_released_open_weights</guid><pubDate>Fri, 31 Jul 2026 07:00:00 +0000</pubDate><description>MiniMax launched H3, a multimodal model that generates up to 15 seconds of 2K video with native stereo audio, plus video-to-video motion transfer and text/brand rendering aimed at commercial content. On Artificial Analysis it leads video editing and beats ByteDance's Seedance 2.0 in some tasks, but trails Google's Gemini Omni Flash on text-to-video and sits behind both on image-to-video. MiniMax says 2K pricing is under a third of mainstream models' rates and plans to release the weights 'in the coming days' under the MiniMax Community License, which permits free non-commercial use and commercial use for organizations under $20M revenue with attribution.

Why it matters: Open weights have barely touched video generation, which remains closed-source and slow-iterating. If H3's weights actually ship at these prices, it's the first credible open base for teams building video pipelines instead of renting an API.</description></item><item><title>Gemini Robotics ER 2 puts an embodied-reasoning brain behind the API</title><link>https://deepmind.google/blog/gemini-robotics-er-2-powering-robotics-with-video-understanding-task-orchestration-and-multi-robot-collaboration</link><guid isPermaLink="false">2026-07-31:multimodal:https://deepmind.google/blog/gemini-robotics-er-2-powering-robotics-with-video-understanding-task-orchestration-and-multi-robot-collaboration</guid><pubDate>Fri, 31 Jul 2026 07:00:00 +0000</pubDate><description>Google DeepMind released Gemini Robotics ER 2, an 'embodied reasoning' model that plans multi-step physical tasks, tracks progress from continuous video, and hands motor execution to any lower-level vision-language-action model while calling tools like Search. It's available now via the Gemini API and AI Studio, integrated with the Gemini Live API for low-latency streaming, and adds multi-robot collaboration. DeepMind reports 57.4% accuracy on progress classification and 91.3% on moment-finding at sub-second latency, and claims one checkpoint can drive different hardware, from Boston Dynamics' Spot to humanoid arms.

Why it matters: The pitch is a general planning layer you can point at whatever robot and VLA you already run, exposed through the same Gemini API developers use for text. It moves robotics tooling from bespoke demos toward something you can actually call.</description></item><item><title>Inflect v2 packs complete TTS into under 4M parameters</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1v5ve6v/i_released_inflect_v2_two_ultratiny_complete_tts</link><guid isPermaLink="false">2026-07-25:multimodal:https://www.reddit.com/r/LocalLLaMA/comments/1v5ve6v/i_released_inflect_v2_two_ultratiny_complete_tts</guid><pubDate>Sat, 25 Jul 2026 07:00:00 +0000</pubDate><description>An independent developer released Inflect v2, two fully local text-to-speech models: Nano at 3.96M parameters (16MB FP32) and Micro at 9.36M. Both include text processing, timing, generation and vocoder — text in, 24kHz speech out, no external vocoder or API. Reported metrics: Micro hits 4.395 UTMOS22 with 3.99% semantic WER at 6.28x real-time on CPU; Nano runs 10.72x real-time. English-only, single fixed voice, no cloning.

Why it matters: A genuinely usable neural TTS stack this small reopens on-device, offline voice for constrained hardware where multi-billion-parameter systems can't go.</description></item><item><title>Black Forest Labs' FLUX 3 fuses video, audio, and robot control into one model</title><link>https://bfl.ai/blog/flux-3</link><guid isPermaLink="false">2026-07-24:multimodal:https://bfl.ai/blog/flux-3</guid><pubDate>Fri, 24 Jul 2026 07:00:00 +0000</pubDate><description>FLUX 3 is a multimodal foundation model that jointly trains on image, video, and audio, built on BFL's Self-Flow method. It generates video with native audio up to 20 seconds, plus text-to-video, image-to-video, video-to-video, keyframe transitions, and agentic clip chaining. In BFL's own preliminary preference tests on 10-second 720p clips it beat Luma Ray 3.2 (93%), Runway Gen-4.5 (77%), and Grok Imagine (69%), but only tied Seedance 2.0 and Gemini Omni Flash at ~52% each; no independent tests exist yet. A spinoff, FLUX-mimic, uses the video backbone as a video-action model for dexterous robotics and is being tested on production tasks at Audi. FLUX 3 Video is in early access; an open-weight backbone called FLUX 3 Dev and a FLUX 3 Image release are slated for the coming weeks.

Why it matters: An independent, open-weights-friendly European lab claiming near-SOTA video+audio and extending the same world model into robot control is a real shot across the bow of both the closed video labs and the VLA robotics crowd.</description></item><item><title>Swiss Apertus 1.5 ships fully open 8B and 70B models with multimodal input and 262K context</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1v539p8/swissaiapertusv15_70b8b</link><guid isPermaLink="false">2026-07-24:multimodal:https://www.reddit.com/r/LocalLLaMA/comments/1v539p8/swissaiapertusv15_70b8b</guid><pubDate>Fri, 24 Jul 2026 07:00:00 +0000</pubDate><description>The swiss-ai team released Apertus 1.5 in 8B and 70B sizes, extending Apertus 1.0 via continued pretraining that added a multimodal mix of 4T tokens (8B) and 2T tokens (70B). The models now accept image, audio, and text input, add an optional thinking mode, and support 262,144-token context, a fourfold increase over 1.0. Post-training improves instruction following and tool use, and the release keeps the fully-open stance: open weights, open training data, and full recipes, with opt-out consent respected retroactively. Architecture is unchanged, a decoder-only transformer with xIELU activations trained with AdEMAMix; a technical report with benchmarks and intermediate checkpoints is promised in the coming weeks.

Why it matters: Truly open data plus weights and recipes remains rare, and a reproducible multimodal model at this scale is a better base for research than the open-weights-only norm.</description></item><item><title>Microsoft's Fara1.5 is a vision-only browser agent, fine-tuned from Qwen</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1v3ny84/microsoftfara1527b_hugging_face</link><guid isPermaLink="false">2026-07-23:multimodal:https://www.reddit.com/r/LocalLLaMA/comments/1v3ny84/microsoftfara1527b_hugging_face</guid><pubDate>Thu, 23 Jul 2026 07:00:00 +0000</pubDate><description>Microsoft Research released Fara1.5, a computer-use agent family (4B, 9B, 27B) that drives web browsers from screenshots alone — no DOM or accessibility tree — emitting click, type, scroll, visit-URL and web-search tool calls with pixel-coordinate arguments. The 27B is supervised fine-tuned from Alibaba's Qwen3.5-27B on trajectories synthesized and verified by Microsoft's FaraGen pipeline, and is designed to deploy with MagenticLite. Microsoft explicitly flags prompt injection embedded in page content, compounding multi-step errors, and hallucinated page state as known limitations.

Why it matters: A capable open-weight CUA that grounds on pixels doubles as a grounding model for other agents — though Microsoft building it atop a Chinese base model is its own quiet commentary on the American open-weights gap.</description></item><item><title>Robotics teams ditch the robot to fix the data bottleneck</title><link>https://the-decoder.com/xiaomi-robotics-1-shows-that-more-data-beats-bigger-models-when-training-robots-to-move</link><guid isPermaLink="false">2026-07-21:multimodal:https://the-decoder.com/xiaomi-robotics-1-shows-that-more-data-beats-bigger-models-when-training-robots-to-move</guid><pubDate>Tue, 21 Jul 2026 07:00:00 +0000</pubDate><description>Xiaomi-Robotics-1 and Hugging Face's Grabette independently attack robot learning's data scarcity the same way: handheld grippers with cameras that a human waves around to record 6-DoF manipulation demos, no robot or teleop rig required. Xiaomi collected over 100,000 hours, auto-labeled it with an LLM in about two weeks, and found more data beats bigger models, with unfamiliar-environment success climbing from ~25% to ~75% as data scaled, beating Physical Intelligence's pi baseline. Grabette is fully open (Raspberry Pi, off-the-shelf OAK-D depth camera, LeRobot format) and pitched as the seed for a shared community dataset; both projects promise code and weights.

Why it matters: If a gripper of commodity parts and a phone-grade camera can generate training data, the VLA data moat weakens and genuinely open robotics datasets start to look feasible.</description></item><item><title>MiniCPM goes embodied with open-source VLA and tracking models</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1v1gcok/minicpmrobot_model_series_minicpmrobotmanip</link><guid isPermaLink="false">2026-07-20:multimodal:https://www.reddit.com/r/LocalLLaMA/comments/1v1gcok/minicpmrobot_model_series_minicpmrobotmanip</guid><pubDate>Mon, 20 Jul 2026 07:00:00 +0000</pubDate><description>OpenBMB open-sourced MiniCPM-Robot, its first embodied-AI series: MiniCPM-RobotManip, a 1.5B general-purpose vision-language-action model for robotic manipulation, and MiniCPM-RobotTrack, a 0.5B model for real-world target tracking. The release ships alongside PhyAI, an inference framework built for embodied models, with weights on Hugging Face.

Why it matters: Sub-2B open VLA models that target real robot hardware push embodied AI toward hobbyist and edge budgets, and give developers a concrete open baseline to fine-tune against instead of closed robotics stacks.</description></item><item><title>DeepMind repurposes a video generator as a computer-vision backbone</title><link>https://the-decoder.com/google-deepmind-argues-video-generators-already-contain-the-world-models-computer-vision-has-been-missing</link><guid isPermaLink="false">2026-07-19:multimodal:https://the-decoder.com/google-deepmind-argues-video-generators-already-contain-the-world-models-computer-vision-has-been-missing</guid><pubDate>Sun, 19 Jul 2026 07:00:00 +0000</pubDate><description>GenCeption takes Alibaba's open-source Wan2.1 video model and, with a one-forward-pass modification, performs depth estimation, segmentation, surface normals and 3D pose from a text prompt. Trained mostly on 7,500 synthetic videos, 7 to 500 times less data than rivals, it matches or beats specialists such as DepthAnything 3 and, on language-guided segmentation, Meta's SAM 3 combined with Gemini 3.5 Flash. It also generalizes to real footage and unseen categories like animals.

Why it matters: A concrete data point that generative video models already carry reusable spatial world models, reviving the pixel-prediction-versus-JEPA debate, though 6-to-10-second-per-clip inference keeps it out of production for now.</description></item><item><title>RadLE 2.0 finds radiology models confidently wrong</title><link>https://the-decoder.com/ai-chatbots-reading-x-rays-can-be-dangerously-confident-even-when-theyre-wrong</link><guid isPermaLink="false">2026-07-19:multimodal:https://the-decoder.com/ai-chatbots-reading-x-rays-can-be-dangerously-confident-even-when-theyre-wrong</guid><pubDate>Sun, 19 Jul 2026 07:00:00 +0000</pubDate><description>Ashoka University's RadLE 2.0 benchmark scored 16 models on 200 radiology cases, rewarding calibrated confidence, penalizing overconfident errors and letting models say I don't know. Radiologists scored 988.7 out of 2,000; the best model managed 758. Claude Fable 5 led on safe and reliable answers, Gemini 3 Pro had the highest raw accuracy, and Meta's Muse Spark 1.1 was best at deferring to a human. Open-weight and medical-tuned models tried to answer nearly every case and were often wrong with high confidence.

Why it matters: For anyone shipping AI into high-stakes decisions, the metric that matters is calibration, not raw accuracy. Models that never abstain are the dangerous ones.</description></item><item><title>NVIDIA's Nemotron 3 Embed 8B tops the RTEB retrieval leaderboard</title><link>https://huggingface.co/blog/nvidia/nemotron-3-embed-wins-rteb</link><guid isPermaLink="false">2026-07-17:multimodal:https://huggingface.co/blog/nvidia/nemotron-3-embed-wins-rteb</guid><pubDate>Fri, 17 Jul 2026 07:00:00 +0000</pubDate><description>NVIDIA released Nemotron 3 Embed, a family of open-weight embedding models with open datasets and training recipes. The flagship 8B (BF16) ranks #1 on the RTEB multilingual leaderboard at 78.5% and 75.5% on MMTEB Retrieval, with 1B BF16 and NVFP4 variants aimed at production; the NVFP4 build claims up to 2x BF16 throughput on Blackwell while retaining 99%+ of retrieval accuracy. All ship day-0 on Hugging Face with a 32k context window, vLLM support, and an optimized NIM microservice, and NVIDIA argues better retrieval cuts downstream agent token costs by returning relevant evidence earlier.

Why it matters: Retrieval quality is the cheapest lever for agent reliability and cost, and an open, fine-tunable embedding model at the top of RTEB gives teams a self-hostable alternative to provider-bundled search.</description></item><item><title>Thinking Machines ships Inkling, a 975B open-weights MoE that leads US labs but trails China</title><link>https://thinkingmachines.ai/news/introducing-inkling</link><guid isPermaLink="false">2026-07-16:multimodal:https://thinkingmachines.ai/news/introducing-inkling</guid><pubDate>Thu, 16 Jul 2026 07:00:00 +0000</pubDate><description>Mira Murati's Thinking Machines released Inkling, its first model: an Apache 2.0 Mixture-of-Experts transformer with 975B total / 41B active parameters, 1M-token context, and native text/image/audio input, pretrained on 45T tokens. Artificial Analysis scores it 41 on its Intelligence Index — the top US open-weights model, ahead of Nemotron 3 Ultra (38) — but it lags GLM-5.2, Kimi K2.6 and DeepSeek v4 on several fronts and posts a rough 63% hallucination rate. Architecturally it drops RoPE for relative positional embeddings and adds short convolutions; a 276B-A12B Inkling-Small preview matches it on some benchmarks. It's on Hugging Face and fine-tunable on Tinker today.

Why it matters: It's the strongest US-origin open-weight release so far and a deliberate bet on customization over leaderboard-maxing — but with post-training bootstrapped from Kimi K2.5, the 'not distilled' purity claims don't hold, and it still trails the Chinese open frontier.</description></item><item><title>Google Images turns 25, gets a Pinterest redesign and in-search image gen</title><link>https://blog.google/products-and-platforms/products/search/google-images-25th-anniversary</link><guid isPermaLink="false">2026-07-15:multimodal:https://blog.google/products-and-platforms/products/search/google-images-25th-anniversary</guid><pubDate>Wed, 15 Jul 2026 07:00:00 +0000</pubDate><description>On Google Images' 25th anniversary, Google is rebuilding it into a browsable, real-time 'For You' gallery with savable collections — a clear play for Pinterest's discovery-and-time-on-site turf. It's also adding image generation directly in AI Overviews using its Nano Banana model, so users can create a visual from a text prompt without leaving Search. Both roll out over the coming weeks, starting on US English desktop.

Why it matters: Folding generation into Search is Google's move to keep image-creation traffic inside its ad ecosystem instead of leaking to ChatGPT — and Nano Banana is now the default engine behind it.</description></item><item><title>audio.cpp 0.3: Supertonic 3 hits 200x realtime TTS on a 5090</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1uwpvt9/audiocpp_10_hours_of_audio_generated_in_3_minutes</link><guid isPermaLink="false">2026-07-15:multimodal:https://www.reddit.com/r/LocalLLaMA/comments/1uwpvt9/audiocpp_10_hours_of_audio_generated_in_3_minutes</guid><pubDate>Wed, 15 Jul 2026 07:00:00 +0000</pubDate><description>The GGML/C++ audio.cpp project shipped release 0.3 with five new TTS models: Supertonic 3, MOSS-TTS-Local, MOSS-TTS-Nano, IndexTTS2, and Irodori-TTS. Supertonic 3 reportedly hits 200x+ realtime on an RTX 5090, 6x+ on CPU, and ~47ms TTFT in CUDA streaming — the demo generated ~10 hours of audiobook audio in about 3 minutes. Because the reference implementation was ONNX and offloaded nodes to CPU, the reverse-engineered C++/safetensors path is markedly faster on GPU; IndexTTS2 longform is 5.65x faster than Python. GGUF support is rolling out model by model.

Why it matters: Local TTS at hundreds of times realtime with sub-50ms latency makes fully on-device voice agents and bulk narration practical without an API bill.</description></item><item><title>Wan-Dancer breaks the 20-second wall for music-to-dance video</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1uvdaq7/wandancer_a_hierarchical_framework_for</link><guid isPermaLink="false">2026-07-14:multimodal:https://www.reddit.com/r/LocalLLaMA/comments/1uvdaq7/wandancer_a_hierarchical_framework_for</guid><pubDate>Tue, 14 Jul 2026 07:00:00 +0000</pubDate><description>Alibaba's HumanAIGC released Wan-Dancer-14B (weights and inference code), a hierarchical framework that generates 720p/30fps dance videos exceeding a minute directly from music. It decouples global keyframe planning from local refinement and uses time-mapped RoPE embeddings plus an optical-flow loss to fight the temporal drift and identity inconsistency that break diffusion models past ~20 seconds, claiming SOTA across five dance genres.

Why it matters: Minute-scale temporal coherence is the actual hard problem in video generation; shipping open weights means the SOTA claim is testable today rather than a demo reel.</description></item><item><title>Google's SensorFM: one foundation model for wearable sensor data</title><link>https://the-decoder.com/sensorfm</link><guid isPermaLink="false">2026-07-13:multimodal:https://the-decoder.com/sensorfm</guid><pubDate>Mon, 13 Jul 2026 07:00:00 +0000</pubDate><description>Google Research unveiled SensorFM, a foundation model pretrained self-supervised on over a trillion minutes of unlabeled Fitbit and Pixel Watch data from five million people across 100+ countries. It processes 34 features from five sensor types (PPG, acceleration, skin conductance and temperature, altitude) and beat supervised baselines with hand-crafted features on 34 of 35 downstream health tasks. Performance scaled cleanly with model and data size, from ~100K to 100M parameters. It remains research-only, aggregated to minute-level data, and tested only on Google's own devices.

Why it matters: It's the wearables version of the 'one big pretrained model replaces many task-specific ones' pattern, and a signal for where personal-health agents get their context. Note the caveats: no raw signals, self-reported labels, and no shipping plans.</description></item><item><title>Moondream 3.1 ships a 9B-A2B MoE vision model</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1uunqcz/moondream319ba2b</link><guid isPermaLink="false">2026-07-13:multimodal:https://www.reddit.com/r/LocalLLaMA/comments/1uunqcz/moondream319ba2b</guid><pubDate>Mon, 13 Jul 2026 07:00:00 +0000</pubDate><description>Moondream 3.1 is a vision-language model with a mixture-of-experts architecture: 9B total parameters, 2B active. It advertises query, detect, point, and caption skills, all returning structured output natively, while staying cheap to deploy. It's pitched as state-of-the-art visual reasoning and detection at small active-parameter cost.

Why it matters: A 2B-active MoE VLM with native structured detection output is a practical building block for local vision pipelines that need bounding boxes and points, not just captions.</description></item><item><title>BAAI's Orca world model matches robot controllers without ever seeing an action label</title><link>https://the-decoder.com/chinas-orca-world-model-matches-specialized-robotics-systems-without-ever-seeing-a-single-action-label</link><guid isPermaLink="false">2026-07-11:multimodal:https://the-decoder.com/chinas-orca-world-model-matches-specialized-robotics-systems-without-ever-seeing-a-single-action-label</guid><pubDate>Sat, 11 Jul 2026 07:00:00 +0000</pubDate><description>Beijing Academy of AI released Orca, a 'world foundation model' that predicts the next abstract world state rather than the next token, frame, or action. Built on a frozen Qwen3.5 core with swappable output heads (text via Qwen, images via Stable Diffusion 3.5, a from-scratch 'Action Expert' for control), the 4B version tops small VLMs on text benchmarks and beats FLUX.2 on image prediction. On five two-armed manipulation tasks it matches π0.5 despite its base model never seeing action data during pre-training — control was learned from just 200 recordings per task.

Why it matters: If a general world model can be fine-tuned into a competent robot controller from a couple hundred demos, it directly attacks robotics' labeled-action data shortage — the constraint that's held embodied AI back.</description></item><item><title>OpenAI's GPT-Live listens and speaks at the same time, offloads reasoning to GPT-5.5</title><link>https://www.reuters.com/business/openai-launches-gpt-live-voice-models-that-listen-speak-simultaneously-2026-07-08</link><guid isPermaLink="false">2026-07-09:multimodal:https://www.reuters.com/business/openai-launches-gpt-live-voice-models-that-listen-speak-simultaneously-2026-07-08</guid><pubDate>Thu, 09 Jul 2026 07:00:00 +0000</pubDate><description>OpenAI released GPT-Live-1 and GPT-Live-1 mini, full-duplex voice models that listen and speak simultaneously, handle interruptions, and use filler words like 'mhmm.' The mini replaces Advanced Voice Mode by default for free users. Crucially, hard queries are delegated to GPT-5.5 in the background while the conversation continues, closing the old intelligence gap: GPQA accuracy rises from 45.3% to 84.2% and BrowseComp from 0.7% to 75.2%. API access is coming soon via a signup form.

Why it matters: The background-delegation architecture is the real trick — it decouples conversational latency from frontier reasoning, and an API would let developers build voice agents that don't feel a generation behind text.</description></item><item><title>Kyutai's Pocket TTS clones a voice from 5s on CPU, MIT-licensed</title><link>https://www.reddit.com/r/LocalLLaMA/comments/1up07mk/kyutais_pocket_tts_clones_a_voice_from_5_seconds</link><guid isPermaLink="false">2026-07-07:multimodal:https://www.reddit.com/r/LocalLLaMA/comments/1up07mk/kyutais_pocket_tts_clones_a_voice_from_5_seconds</guid><pubDate>Tue, 07 Jul 2026 07:00:00 +0000</pubDate><description>Kyutai's Pocket TTS is a ~100M-parameter streaming language model that generates audio tokens over the Mimi neural codec and does zero-shot voice cloning from a 5-second reference clip—on CPU, no GPU, no fine-tuning. In a 180-run head-to-head against Kokoro 82M, Supertonic 3 and Inflect-Nano on a 4-core Xeon, it was the slowest config (RTF ~0.71, UTMOS 4.10) but the only model in the field capable of user-supplied voice cloning; latency stays flat across text lengths because it streams token by token. Install is a plain pip install pocket-tts with no CUDA build.

Why it matters: The MIT license plus CPU-only cloning makes it the first genuinely commercial-friendly option for arbitrary-voice TTS on commodity hardware—a category of one against Apache and OpenRAIL competitors.</description></item><item><title>Baidu's Unlimited OCR keeps the KV cache flat across dozens of pages</title><link>https://the-decoder.com/baidus-unlimited-ocr-processes-dozens-of-document-pages-in-one-pass-by-treating-memory-like-human-forgetting</link><guid isPermaLink="false">2026-07-06:multimodal:https://the-decoder.com/baidus-unlimited-ocr-processes-dozens-of-document-pages-in-one-pass-by-treating-memory-like-human-forgetting</guid><pubDate>Mon, 06 Jul 2026 07:00:00 +0000</pubDate><description>Baidu built on the open DeepSeek OCR model with Reference Sliding Window Attention (R-SWA): generated tokens attend to all visual/prompt tokens but only the last 128 output tokens, keeping the KV cache constant instead of growing with document length. The 3B MoE (~500M active) processes 40+ pages in a single pass at edit distance below 0.11, scores 93% on OmniDocBench v1.5 (six points over the DeepSeek OCR baseline), and runs ~12.7% faster in Base mode. Code and weights are on GitHub/Hugging Face with vLLM and SGLang support.

Why it matters: Constant-memory long-document OCR is directly useful, and the underlying trick — cramming text into cheap image tokens — is the same lever people are pulling to extend context windows and cut token bills.</description></item><item><title>Google DeepMind buys into A24 for filmmaking-tools research</title><link>https://deepmind.google/blog/google-deepmind-and-a24-announce-first-of-its-kind-research-partnership</link><guid isPermaLink="false">2026-07-04:multimodal:https://deepmind.google/blog/google-deepmind-and-a24-announce-first-of-its-kind-research-partnership</guid><pubDate>Sat, 04 Jul 2026 07:00:00 +0000</pubDate><description>Google DeepMind and studio A24 announced a multi-project research partnership (reported at $75M, including a Google investment) to develop new filmmaking workflows and tools via A24 Labs, anchored on systems like Gemini and Veo. Coverage frames it as DeepMind borrowing A24's cultural credibility to make its AI ambitions 'feel cooler and more inevitable' — and notes a chunk of Hollywood is quietly rooting for the deal to collapse.

Why it matters: It's a bet that generative video's adoption problem is taste and trust, not just model quality — and a test of whether a prestige brand can partner with a hyperscaler without diluting itself.</description></item></channel></rss>
