daily
Jul 26, 2026

AI Daily — 2026-07-26

English 中文

OpenAI's Altman previews its most powerful AI in Washington as ChatGPT expands with trip-planning and emails, while Open-Weight 4B models achieve O3-l


Covering 43 AI news items

🔥 Top Stories

1. OpenAI’s Altman to Washington to preview its most powerful AI yet

Sam Altman heads to Washington to preview OpenAI’s most powerful AI and pushes for swift regulatory approval. The post claims the model just hacked a real company and speculates about GPT-6, describing capabilities like coordinating swarms of agents and long-horizon reasoning for government and business. Source-twitter

2. ChatGPT proves versatile, builds trip-planning site and emails

A user tweets that ChatGPT used their chat history to generate weekend-trip ideas, plan three options, build a full-stack coordination site for eight friends, and draft a Gmail email when a group decision is reached. The demo highlights AI-assisted planning, development, and automation in a single workflow. Source-twitter

3. Open-Weight 4B Models Reach O3-Level Swedish Medical Q&A

Open-weight 4B models are nearing O3-level performance on Swedish medical licensing questions. In MedQA-SWE, o3 achieves 88% accuracy in 2025 (vs GPT-4’s 84% in 2024); with post-training, MedGemma-1.5-4B reaches 60% on the final year, while Gemma4-E4B and Qwen3.5-4B reach 77% untrained and about 87% with reasoning, though loops can occur. The discussion cites an S-GRPO early-exit approach and provides a GitHub implementation. Source-reddit

LLMs

  • Frontier LLMs Achieve Near-Perfect IMO 2026 Scores; Web Apps Lag — Frontier models Sol and Fable achieved perfect or near-perfect scores on IMO 2026, largely irrespective of harnessing. Webapp performance lagged, but improved with Claude Code and further boosted by AutoFyn, a multi-agent harness developed by the authors; GLM’s performance was comparable to Sonnet without harness and improved with AutoFyn. Numerical scores are provided in the attached paper: https://preview.redd.it/fy4ayale5nfh1.png?width=2155&format= Source-reddit

LLM

  • Open-source SDLC harness beats Claude Code via repo learning — AutoDev Studio is an open-source AI coding agent that builds a persistent, local knowledge base from the repository to reuse localization across tasks. In benchmarks across six well-localized tasks on repos up to ~82k LOC, it was 7%-75% cheaper than a cold Claude Code run (e.g., $6.83 vs ~$1.70). The workflow includes clarifying questions, coding on an isolated branch, QA tests, cross-model reviews, a bounded revise loop, and opening a real GitHub PR; details are in the README. Source-reddit
  • Ollama Signs Satya Nadella Letter on Open-Weight Models — Ollama signs Satya Nadella’s letter advocating open models accessible to every developer to unlock the next frontier for America and the globe. Nadella’s message argues that open-weight models are essential to a healthy AI ecosystem and outlines a path to strengthen American competitiveness and expand economic opportunity while safeguarding national security. Source-twitter
  • Claude Code Interface Enables Multi-Model Coding with Fugu-Ultra v1.1 — SakanaAI Labs announces a Claude-compatible interface for Fugu-Ultra v1.1, enabling orchestration of a diverse pool of frontier models inside the coding workflow. Instead of relying on a single model to write, debug, and execute code, developers can manage multiple models from the terminal, expanding collaborative AI coding capabilities. Source-twitter
  • ChatGPT Work Overtakes Codex in Active Users — A tweet claims ChatGPT Work now has more active users than Codex. The post expresses skepticism about the use case for ChatGPT Work, but acknowledges the growth in adoption. Source-twitter
  • obra’s Superpowers: AI coding agents framework and tooling — obra’s Superpowers offers a complete software development methodology for coding agents, built on composable skills and initial guidance. The project lists a Quickstart with tools like Claude Code, Codex, Gemini CLI, and others, and it announces a full-time community engineer role with details on the GitHub page. It emphasizes a conversational workflow where the agent clarifies the user’s goal and teases a spec before writing code. Source-github
  • GPT-5.5 scores 10.6% on ActiveVision; humans 96.1% — ActiveVision is a 17-task, 3-category benchmark designed to force repeated visual perception. GPT-5.5 achieves 10.6% accuracy (zero on 11 tasks); Claude Fable 5 scores 3.5%; humans average 96.1%. The study highlights current frontier models’ struggles in this regime and their inability to patch it by self-modifying code. Source-reddit

Optimization

  • SkewAdam Cuts MoE Optimizer State by 97% on 40GB GPUs — SkewAdam introduces a tiered state allocation to dramatically reduce optimizer memory in Mixture-of-Experts training. It lowers optimizer state memory from 50.6 GB to 1.29 GB and peak training memory from 81.4 GB to 31.3 GB, enabling a 6.78B MoE to fit on a single 40GB GPU without sacrificing convergence or router performance. The method assigns precision by parameter group: backbone (5%) with momentum plus factored 2nd moment; experts (95%) with factored 2nd moment only; router (<0.01%) with exact 2nd moment. Source-reddit

Open Source

  • Need for AI competition and open weights — A tweet argues that AI requires competition and open weights to prevent a single firm from becoming the moral arbiter of acceptable speech. It cites Anthropic as an example and notes Grok and Kimi K as models that operate without such restrictions, calling the central idea ridiculous. Source-twitter
  • Open-Weight Models Matter as Nanbeige 4.2 Highlights Looped Depth — Open-source/open-weight AI models are highlighted as essential for verification, transparency, and running AI on personal hardware outside closed labs. The post notes upcoming Kimi K3 and Ling 3.0 weights, and surveys several new open-weight releases, including Nanbeige 4.2 3B which uses looped depth sharing to double the transformer blocks without duplicating weights. It frames the week as a notable roundup of architecture innovations in open-weight AI. Source-twitter
  • mattpocock/skills: Open-source agent skills for real engineers — mattpocock/skills offers small, adaptable agent skills designed for real engineering work, emphasizing developer control over heavy-process frameworks. The open-source project is model-agnostic and installable via npx, with a dedicated setup script for agents. It promotes hands-on engineering over hype and references approaches like GSD, BMAD, and Spec-Kit while inviting readers to join the author’s newsletter for updates. Source-github
  • Open-weight AI is having its Kubernetes moment — Open-weight AI models are gaining traction as Kubernetes becomes the preferred deployment platform. The piece highlights the shift toward open weights and tooling that enable scalable inference, orchestration, and ecosystem interoperability on Kubernetes. This trend could accelerate open-weight adoption across AI workloads. Source-hackernews

AI Policy

  • OpenAI Signs Open Weights AI Leadership Letter — A Twitter thread notes that competition benefits the AI ecosystem and that serving models at scale is hard. It reports that OpenAI signed the Open Weights and American AI Leadership letter, and that OpenClaw signed Microsoft’s Open Weights and American AI Leadership letter, while Ant remains silent. The Open Weights initiative is framed as protecting user choice and enabling people to run, study, and build AI on their own terms. Source-twitter
  • What is happening to jobs? Separating AI hype from reality — A Stanford SEIPR policy brief analyzes AI’s impact on employment, separating hype from actual effects. It argues that AI-related disruption is nuanced, driven by productivity gains and shifting skill demands rather than instant, widespread job loss. The document calls for evidence-based policymaking to support workers and firms as AI adoption evolves. Source-hackernews

AI

  • ChatGPT Becomes Your Personal AGI for Everyday Tasks — A social post claims ChatGPT can act as a personal AGI, letting users “work” for them by prompting the AI to handle everyday tasks from a mobile device. Cited examples include negotiating internet bills, unsubscribing from newsletters, and finding deals, with the author noting it performs many tasks daily and remains surprised by its capabilities. Source-twitter
  • Running a 28.9M parameter LLM on an $8 microcontroller — An experiment demonstrates running a 28.9M parameter LLM on an $8 microcontroller, pushing edge AI capabilities on extremely constrained hardware. The project, esp32-ai by slvDev, shows on-device inference for a mid-sized model on a low-cost ESP32-class device. The effort has garnered discussion on Hacker News. Source-hackernews
  • Compiler turns computation graphs into vanilla transformer weights with no training — A researcher built a compiler that takes a Python computation graph and outputs the weights of a vanilla transformer that can execute the graph, with zero training in the pipeline. The resulting Phi-3 architecture checkpoint loads natively in Hugging Face with no custom code or trust_remote_code. The project includes a write-up and a twelve-example repo, and cites RASP and Tracr as related approaches. Source-reddit

Theoretical AI

  • NeurIPS 2026 Theory Track: Initial Review Scores Shared — A Reddit post asks for early review distributions for NeurIPS 2026 Main Track theory papers. The author reports their paper received 4/3/3 with confidence 3/3/3, noting theory papers may receive conservative early scores, and invites others with theory submissions to share their initial scores to identify patterns. Source-reddit

AI Safety

  • NeurIPS 2026 reviews flagged for prompt-injection concerns — Reddit user reports that a NeurIPS 2026 submission downloaded from OpenReview triggered a GPT prompt-injection warning, which the user says they did not insert. They compared versions and suspect NeurIPS added the injection. The post invites others to share similar experiences and suggests checking reviewer copies for unusually formulaic phrasing to spot potential LLM-generated reviews. Source-reddit

⚡ Quick Bites

  • Anthropic and Dario Amodei in Open-Source Lobbying Debate — A tweet accuses Anthropic and Dario Amodei of lobbying against open-source AI, claiming their supporters attack companies backing open-source efforts. It also notes Julian Schrittwieser praising Jensen Huang’s embrace of open source and anticipation of CUDA and GPU driver open-source releases. Source-twitter
  • GLM 5.2 NVFP4: 100M tokens locally for $1 — User Alec Fong claims he ran 100 million tokens through GLM 5.2 NVFP4 locally at about $1 in inference cost. The post implies on-device AI inference is becoming nearly free, highlighting efficiency gains for local LLM deployment. Source-twitter
  • Feedback Requested: What should the next Gemma models offer? — An X (Twitter) post by osanseviero invites feedback on the next Gemma models, asking what capabilities should be added and why. The tweet seeks input from the community to guide future Gemma development and prioritization of features. Source-twitter
  • Terence Tao: Mathematics in the Age of AI — Terence Tao delivers slides for ICM 2026 exploring how artificial intelligence shapes mathematics, including implications for proofs, intuition, and collaboration between humans and AI. The talk surveys opportunities and challenges at the intersection of AI and mathematical practice. Source-hackernews
  • Cloudflare’s new AI traffic options for customers — Cloudflare announces new options for handling AI-related traffic for its customers. The blog post explains how these options help manage AI workloads, routing, and performance on Cloudflare’s platform. Source-hackernews
  • AI Mania Eviscerates Global Decision-Making — A critical take argues that rampant AI hype is reshaping and potentially degrading global decision-making. It warns against overreliance on AI tools and opaque optimization, urging more cautious, human-led governance. The piece is linked through Daring Fireball and discussed on Hacker News, highlighting the debate around AI influence. Source-hackernews
  • Claude 5 Context Engineering Rules for Generation Models — The article outlines updated context engineering guidelines for Claude 5 generation models, detailing recommended prompt structuring, context window management, and retrieval strategies to improve performance and safety. It discusses practical implications for developers and users applying Claude 5 in real-world tasks. Source-hackernews
  • Debian Considers Three Proposals for LLM Usage — Debian is voting on three proposals about how to use large language models (LLMs) within the Debian ecosystem. The vote_002 page outlines the proposals, and public discussion on Hacker News shows notable engagement. The outcome will shape Debian’s policy and tooling for AI usage. Source-hackernews
  • Politician Reads AI Prompt During Assembly — During an assembly, a politician read an AI prompt on stage, drawing attention to how AI prompts influence political discourse. The moment underscores ongoing debates about AI governance, transparency, and the role of AI tools in public deliberation. The clip gained attention on YouTube and was discussed on Hacker News. Source-hackernews
  • AI jobs apocalypse unlikely in near term — Guardian analysis argues that a sudden AI-driven jobs apocalypse is unlikely in the near term. It suggests automation will reshape work gradually rather than wipe out human labor, with productivity gains and new opportunities offsetting displacement. The piece cites studies and expert opinions that emphasize policy responses like retraining and safety nets. Source-hackernews
  • YOLO26n Inference Rewritten in ARM64 Assembly on Raspberry Pi — A developer implemented YOLO26n inference from scratch using ARM64 Assembly and C, without frameworks, to study low-level neural network engines and edge AI optimization on Raspberry Pi 4. The project features NEON SIMD, Winograd convolution, optimized GEMM kernels, cache-aware tiling, custom micro-kernels, operator fusion, and attention within the YOLO26 components, with a redesigned memory layout in a custom binary format. While the detector produces correct results, performance gains were below expectations, and feedback is sought. Source-reddit
  • Using AI coding agents with remote GPUs for ML workflows — A software engineer seeks a platform to use AI coding agents (Codex, Claude Code, OpenCode) while running ML code on cloud GPUs. The goal is to code locally with an AI assistant and have the actual ML execution occur on a remote GPU, enabling a seamless develop-debug-iterate loop as if the GPU were attached to the local environment. The user asks for recommendations on tools or platforms that support this workflow. Source-reddit
  • Document Layout Tools: DocLayout, MinerU, Marker, Unlimited-OCR Debate — Reddit user compares document-layout models (DocLayout, Docling, MinerU, Marker, Unlimited-OCR) for journal PDFs. They note Docling is strong but can overfit, MinerU misses items like the corresponding author in the footer and masthead marks, and Unlimited-OCR struggles with styles and logos. They ask what SOTA models are best for PDF text and layout extraction. Source-reddit
  • MCP workflow turns engineering plans into deep-learning implementations — An MCP workflow guides ML engineers from an engineering plan to functioning deep-learning components. It employs Codex to break the plan into blocks, identify relevant papers, extract implementation details, prepare specifications, implement components in dependency order, and record results. Papers are used as supporting references, not to define the project or reproduce specific works. Source-reddit
  • One Encoder, Seven Heads: Lessons From Masked-Loss Multi-Task Classifier — Researchers consolidated seven sequence classifiers into a single multi-head model with a shared mmBERT-small encoder. They masked absent-task losses and added a self-test to ensure gradients for missing tasks are zero, catching two subtle bugs. About 5k synthetic/real multi-task rows aided training, while held-out evaluation across heads yielded high F1 scores in the mid-to-high 0.9s on real data. Source-reddit
  • GPT-Live Enhances Reading Experience, José Ocampo Says — A tweet describes reading a book while GPT-Live is open to chat, ask questions, take notes, and bookmark. José Ocampo calls the experience like reading in 4D, highlighting a more interactive reading workflow powered by AI. Source-twitter
  • Hermes Desktop 18h: Open-Source Harness for Team Collaboration — A tweet discusses Hermes Desktop 18h and asks whether there is a user-friendly harness like Claude cowork that integrates with open-source models and supports company multiplayer modes. It highlights interest in easy-to-use tooling and multi-user collaboration for AI model deployment. Source-twitter
  • AI Superpowers: Focus and Followthrough — An AI-focused newsletter argues that focus and follow-through have become the new superpowers in AI work. It contends that turning research into practical, deployed systems requires disciplined prioritization, rigorous experimentation, and effective productization, rather than chasing hype. The piece emphasizes execution as the key driver of real-world AI impact. Source-hackernews
  • AI Productivity Illusion: Gains Overstated — The article argues that AI tools promise productivity boosts, but real-world gains are often overstated due to integration friction and misaligned incentives. It examines scenarios where AI underdelivers and emphasizes careful implementation and expectation management. The piece also includes reader perspectives from Hacker News. Source-hackernews
  • Understanding GPU Inference Workloads and Sourcing Compute — This post explores how people source compute for GPU inference workloads and highlights pain points in the process. It asks for experiences with online services like runpod and vast.ai, inviting comments or DMs. A short, 2-minute survey link is also requested. Source-reddit
  • Adyen ML Interview: What to Expect From HackerRank Round — A candidate received an ML interview offer from Adyen and will have a live HackerRank coding round. They are unsure whether questions will be DS&A-focused or ML/data-oriented and are seeking others’ experiences and prep tips. Source-reddit

Generated by AI News Agent | 2026-07-26