AI Daily — 2026-09-09
New AI models feature open-source speech generation/editing, real-time full-duplex multimodal interaction, and self-improvement via agentic post-train
Covering 25 AI news items
🔥 Top Stories
1. AuK Open-Source Model Unifies Speech Generation and Editing
AuK is a foundational speech model that handles generation, content editing, enhancement, separation, paralinguistic editing, and acoustic editing through natural-language instructions. Trained on roughly 3.03 billion instruction-audio instances and 1.95 million hours of supervision, it offers an unusually broad interface for speech control. The open-source release is significant for developers building speech assistants, audiobook tools, and audio post-production workflows. Source-huggingface
2. Gander Combines Omni Perception, Full-Duplex Interaction, and Agentic Skills
Gander is an end-to-end model that processes streaming video, speech, and text inputs while supporting natural real-time conversations and agentic workflows. Its full-duplex design means users can interrupt the model at any time, a key requirement for human-like voice and multimodal interfaces. This suggests a path toward more unified assistants that can perceive, converse, and act without switching between separate specialist models. Source-huggingface
3. 27B 1-Bit Model Runs in the Browser at 25–30 tok/s on a 6GB GPU
A solo developer built a WebGPU/WGSL inference engine that runs Bonsai-27B, a 27B-parameter 1-bit model, entirely in the browser at 25–30 tokens/s on a 6GB RTX 3060 Laptop. The model fits in 3.8GB of VRAM and requires no installation or server, showing how far on-device inference has come. This milestone could accelerate private, low-friction deployment of large models on consumer hardware. Source-reddit
📰 Featured
AI Research & Safety
- NeoHorse-1 Explores Self-Improvement via Agentic Post-Training — This family of agent-native models uses a heterogeneous model pool with intelligent routing to enable recursive self-improvement, raising both capability and safety questions for future training cycles. Source-huggingface
- On-Policy Reverse Distillation Boosts Weak-to-Strong Generalization — OPRD lets stronger student models surpass weaker teachers by distilling from their own on-policy outputs, potentially enabling more efficient consolidation across successive model generations. Source-huggingface
Autonomous Driving
- DriveZero: End-to-End Driving System That Learns Beyond Human Demos — DriveZero splits driving into separately pretrained perception and action models before merging them into a unified planner, aiming to exceed the ceiling of human demonstration logs in autonomous driving. Source-huggingface
Agents & Developer Tools
- Superpowers Framework Teaches Coding Agents to Ask Before Coding — This composable skills framework pushes coding agents to clarify user intent and generate a spec before writing code, potentially reducing wasted iterations on misunderstood requirements. Source-github
- Browser Use Lets AI Agents Automate Web Tasks — An open-source Python library and managed cloud service give AI agents browser control, enabling automation of form filling, data extraction, and more complex web workflows. Source-github
- OpenAI Publishes Curated Codex Plugin Examples on GitHub — The repository includes plugin integrations for Figma, Notion, and app development, with manifest examples and companion surfaces to help developers build on Codex. Source-github
Hardware & On-Device Inference
- Apple A20 Pro Chip Features 7-Core GPU, 32-Core Neural Engine — The reported A20 Pro uses a 2nm process, doubled Neural Engine cores, and a 96-bit LPDDR5X memory bus, delivering roughly 50% more memory bandwidth at around 115 GB/s. Source-reddit
⚡ Quick Bites
- NVIDIA Cosmos3 64B Runs Locally via INT4 Quants on Apple Silicon — Reddit users report that NVIDIA’s Cosmos3 64B image model can now run locally on Apple Silicon with INT4 quantization. Source-reddit
- AMD Announces Threadripper Halo Station Workstation for Local AI — AMD’s new workstation is positioned as a serious local AI machine for high-end on-premises workloads. Source-reddit
- AI Audio Model Generates Synths from Text Prompts with Timbre Control — A developer-trained audio model can generate synthesizer sounds from text prompts while offering timbre control. Source-reddit
- ADHD-Friendly Skill Improves Coding Agent Outputs — A GitHub skill package aims to make coding-agent outputs more structured and easier to follow for ADHD users. Source-github
- Karpathy-Inspired CLAUDE.md Improves Claude Code Behavior — A community CLAUDE.md file, inspired by Andrej Karpathy’s workflow preferences, helps guide Claude Code toward cleaner coding behavior. Source-github
- GLM 5.3 Flash Doubles Speed on M3 Ultra via Kernel Fusion — Kernel-fusion optimizations reportedly double GLM 5.3 Flash throughput on Apple M3 Ultra machines. Source-reddit
- Qwen Model Cuts Reasoning Tokens via Noise Reduction — A community GGUF build of Qwen applies noise-reduction techniques to reduce the number of reasoning tokens produced by the model. Source-reddit
- Qwen3-27B Runs at 10–20 TPS on RTX 3060 — Users demonstrate that Qwen3 27B with IQ3_XXS quantization runs at workable speeds on an RTX 3060. Source-reddit
- DeepSeek Soft-Retires V4 Pro Model — Reddit users note that DeepSeek has quietly stepped back from promoting DeepSeek V4 Pro. Source-reddit
- Best Open Source TTS Models for Narration Discussed — The LocalLLaMA community shares current recommendations for open-source TTS models suited to narration. Source-reddit
- Reddit Thread Seeks Best Local Vision Language Models — Users exchange top picks and practical tips for running vision language models locally as of August 2026. Source-reddit
- LM Studio Users Struggle to Download App Amid Bionic Agent Push — Some users report frustration with LM Studio’s download and update process as the project emphasizes a new Bionic Agent feature. Source-reddit
- Don’t Let FOMO Drive Local LLM Buying, Learn with What You Have — Community advice encourages newcomers to experiment with existing local models and hardware before making new purchases. Source-reddit
- User Proposes Clearer Finetune Tags for New Model Releases — A Reddit user suggests that new model releases explicitly indicate when they are finetunes rather than base models. Source-reddit
- Custom Loop Rebuild Drops Temps for 70GB VRAM Server with 2x RTX Titans and 2080ti — A user’s custom loop rebuild reportedly lowers temperatures on a 70GB VRAM setup built from two RTX Titans and a 2080 Ti. Source-reddit
Generated by AI News Agent | 2026-09-09