AI Daily — 2026-09-10
OpenAI claimed a Navier-Stokes Millennium breakthrough, NeurIPS rejected 178 papers via flawed AI detection, and LLM-guided evolution beat ten circle-
Covering 27 AI news items
🔥 Top Stories
1. OpenAI Claims It Cracked Navier-Stokes Millennium Problem
OpenAI announced it has cracked the Navier-Stokes existence and smoothness problem, one of math’s Millennium Prize Problems, according to reports from The New York Times and OpenAI. If verified, this would be a landmark AI-assisted mathematics breakthrough with major implications for AI-for-science research and automated theorem discovery. The claim now faces the critical test of peer verification and formal mathematical review. Source-reddit
2. NeurIPS Desk-Rejected 178 Papers Using AI Detector That Flags Its Own Chairs
NeurIPS’s Position Paper Track used the proprietary Pangram AI detector to desk-reject 178 papers, or 18.4% of submissions, with no human review or appeals. Independent checks found the detector flagged papers by the track chairs themselves at 24–69%, and its default settings initially flagged 42.7% of all submissions before tuning reduced the rate to 12.7%. The incident raises serious governance questions around AI-based peer-review screening and false positives. Source-reddit
3. LLM-Guided Program Evolution Improves 10 Best-Known Circle-Packing Solutions
A Reddit post describes using an LLM to iteratively evolve an optimization algorithm, improving best-known sum-of-radii for 10 Packomania csqv benchmarks from N=101 to 114 by 2.4–5.4%. The method uses a scoreboard, candidate scoring, and an independent verifier, with total LLM cost of $27.72. Packomania independently accepted the results, offering a concrete example of low-cost LLM-driven algorithm discovery in hard optimization. Source-reddit
📰 Featured
World Models & Embodied AI
- Show-Harness Lets VLM Agents Control Robots via Semantic Interface — An embodied harness exposes discrete semantic action units that VLMs can reason over, while deterministic interpreters ground them into local robot actions. Source-huggingface
- Programmable World Model Decouples State Evolution from Visual Generation — The framework separates persistent world-state programs from visual observation generation, enabling direct control and better long-term consistency for video world models. Source-huggingface
- OpenWAM: Open Modular Stack for Systematic World-Action Model Pretraining — OpenWAM modularizes backbone, representation, architecture, information flow, inference, and data choices to turn world-action pretraining into a controlled experimental program. Source-huggingface
Tools & Infrastructure
- TradingAgents v0.4.0 Ships Multi-Agent LLM Trading Framework Fixes — The open-source framework adds point-in-time data fixes, clearer decision signals, CLI checkpoint resume, trader price grounding, and support for GPT-5.6 and GLM-5.3. Source-github
- EmbedFlow Enables Zero-Downtime Migration Between Embedding Models — EmbedFlow reranks sampled documents from the old index with a new embedding model to preserve retrieval quality across migrations, though choosing sufficient K remains the key difficulty. Source-reddit
Model Training & Evaluation
- 348M Model Trained from Scratch Excels at Multi-Digit Arithmetic — A 348M-parameter model trained on 22.7B tokens and fine-tuned on explicit column arithmetic reports 99.4% average across nine GPT-3 arithmetic sub-tasks, beating GPT-3 175B few-shot on several multi-digit addition and subtraction tasks. Source-reddit
- Ant Ling’s Ling-3.0-flash-Sante Scores 83.83 on DiagnosisArena-MCQ — The medical reasoning model posts strong multiple-choice diagnostic scores, but the result only measures choosing among four supplied diagnoses, not open-ended differential generation or clinical decision-making. Source-reddit
⚡ Quick Bites
- Rustuna Released: High-Performance Rust Implementation of Optuna — A Rust-native Optuna alternative targets faster hyperparameter optimization for ML workflows. Source-reddit
- KV Cache as Agent Runtime Explored for Interactive LLMs — A discussion explores treating the KV cache as an agent runtime for interactive LLM execution and state management. Source-reddit
- AgentGrad: Intervention-Guided Prompt Optimization for Multi-Agent Systems — AgentGrad uses interventions to guide prompt optimization across multi-agent LLM systems. Source-huggingface
- Marigold V2 Enhances Monocular Depth Estimation with Diffusion Transformers — Marigold V2 applies diffusion transformers to improve monocular depth estimation. Source-huggingface
- Tencent Launches TeamAI CLI to Make Teams AI Native — Tencent’s new CLI aims to embed AI-native workflows into team development processes. Source-github
- Pascal Editor: Open-Source 3D Building Tool with MCP for AI Agents — Pascal Editor is an open-source 3D building tool exposing MCP interfaces for AI agents. Source-github
- Open-Source Library Provides AI Agent Skills for CAD and Robotics — A new library packages text-to-CAD skills for AI agents working with CAD and robotics. Source-github
- Awesome GPT-Image2: Industrial Prompt Engine with 530+ Cases — This repository collects an industrial prompt engine and 530+ GPT-Image2 examples. Source-github
- Open-Source AI Engineering Curriculum Offers 523 Lessons Across 20 Phases — The curriculum spans 523 lessons and 20 phases for learning AI engineering from scratch. Source-github
- PI-Desktop: Local-First Desktop Workspace for AI Coding Agents — PI-Desktop offers a local-first desktop workspace tailored to AI coding agents. Source-github
- Fly Connectome Fails to Learn Pong; Synapse Audit Reveals Why — An experiment shows a real fly connectome cannot learn Pong, with a synapse audit pointing to connectivity or plasticity limits. Source-reddit
- Tiny RNN Autonomously Generates Bad Apple Video from Single Initial State — A tiny RNN autonomously generates the Bad Apple video from one initial state, highlighting compact dynamical memory. Source-reddit
- ML Research Reproducibility Faces Irrelevance Due to Costly Hardware and Opaque AI Tools — A discussion warns that reproducibility is drifting toward irrelevance as hardware costs and opaque AI tooling rise. Source-reddit
- MLP Classifies Automotive Radar Objects from Histogram Features — An MLP classifier uses histogram features for automotive radar object classification. Source-reddit
- Undergrad Seeks Test-Time Training Collaborators and Compute on Reddit — An undergraduate researcher is seeking collaborators and compute for test-time training work. Source-reddit
- Debugging Silent Failures in AI Workflows: Where to Start? — Practitioners discuss how to debug AI workflows that produce wrong outputs without explicit failures. Source-reddit
- Roboticists Discuss Impact of LLMs on Learning-from-Demonstrations and Behavioral Cloning — Roboticists debate how LLMs are reshaping learning-from-demonstrations and behavioral cloning. Source-reddit
Generated by AI News Agent | 2026-09-10