daily
Aug 18, 2026

AI Daily — 2026-08-18

English 中文

OpenAI paused frontier RL training for safety, Anthropic warned of superintelligence by 2026, and GLM-5.3 launched.


Covering 38 AI news items

🔥 Top Stories

1. OpenAI Pauses Frontier RL Training to Ensure Safety Standards

OpenAI has paused some frontier reinforcement learning runs to validate alignment, security, and monitoring standards before continuing. The move signals that safety is increasingly setting the pace of AI progress, even as the company remains committed to making frontier capabilities widely available. Source-x

2. Anthropic Warns Superintelligence Could Arrive by Late 2026

Anthropic reiterates that model capabilities are advancing rapidly, with open-weight systems becoming competitive with proprietary models, and warns that powerful AI could arrive by late 2026 or early 2027. The accelerating timeline has led some frontier labs to restrict access, highlighting the growing tension between openness and safety. Source-x

3. GLM-5.3 API Launches with Coding, Cybersecurity, Agentic Focus

Z.ai unveiled the GLM-5.3 API, purpose-built for coding, defensive cybersecurity, and long-horizon agentic tasks at the same price as GLM-5.2. Available through the official API and partner model gateways, the release expands competitive options for agentic and security-heavy AI workloads. Source-x

AI Safety & Regulation

  • Anthropic Watermarks AI Text, Revenue Hits $65B, AI Goes Agentic — Anthropic begins watermarking AI-generated text for EU compliance while reaching $65B annualized revenue, with the broader industry shifting toward agentic AI as Nvidia open-sources a physical AI toolkit and Gartner predicts agents in 40% of enterprise apps by 2026. Source-reddit

Industry & Products

  • Claude gains Gmail and Google Drive integration — Claude can now draft and send emails and manage Drive files with user-controlled approval, rolling out across all paid plans. Source-x
  • Harvey II: Smarter AI Agents for Legal Work — Harvey announced Harvey II with Harvey Tenet, its first model trained specifically for legal work, organizing agents around matters and projects with task assignment and review. Source-x

Open Source & Infrastructure

  • Mojo AI Language Released as Open Source — Modular open-sourced Mojo, its high-performance programming language for AI workloads, letting developers freely use and contribute to the toolchain. Source-x
  • StateM Reaches 95.3% Accuracy on Terminal-Bench 2.1 via Harness Scaling — The agent-native runtime achieves 95.3% raw accuracy on Terminal-Bench 2.1 using durable states, phase-local context, and recoverable runbooks without modifying model weights. Source-huggingface

Benchmarks & Evaluation

  • HarnessEval-W Benchmark for Reasoning-Based World Model Evaluation — A new benchmark requires reasoning chains to justify world-model scores, automatically detecting violations of physics, causality, and world state in visual rollouts. Source-huggingface
  • VibeWorlding Framework Trains Multimodal Agents for 3D World Construction — VibeWorlding provides a unified framework for benchmarking and training multimodal agents that construct interactive 3D worlds, testing user intent inference, tool use, and reasoning over textual and visual 3D information. Source-huggingface

⚡ Quick Bites

  • New Paper Introduces Large Discovery Models for Open-Ended Search — A new paper proposes Large Discovery Models for LLM-driven open-ended search and discovery. Source-huggingface
  • New RL Method Improves Multi-Reward Policy Optimization for LLMs — A new reinforcement learning method boosts multi-reward policy optimization for LLM training. Source-huggingface
  • MoneyPrinterTurbo: AI Tool Generates Short Videos from Keywords — MoneyPrinterTurbo is an open-source AI tool that automatically generates short videos from keyword prompts. Source-github
  • Strix: Open-Source AI Pentesting Tool Finds App Vulnerabilities — Strix is an open-source AI pentesting tool that discovers application vulnerabilities. Source-github
  • ai-memory Tool Gives AI Coding Agents Long-Term Memory — The ai-memory tool provides AI coding agents with persistent long-term memory across sessions. Source-github
  • Open-Source Library Offers 817 Cybersecurity Skills for AI Agents — A new open-source library equips AI agents with 817 cybersecurity skills. Source-github
  • llmfit: Open-Source Tool to Match LLMs to Your Hardware — llmfit helps practitioners find the best LLM for their local hardware constraints. Source-github
  • LLM Use Linked to Weaker Critical Thinking, Studies Find — Studies link heavy LLM use with diminished critical-thinking ability. Source-reddit
  • AI Pricing Market Unhinged: From $0.03 to $600 per Million Tokens — A look at the chaotic AI pricing landscape, which ranges from $0.03 to $600 per million tokens. Source-reddit
  • Qwen vs GPT-5.6 vs Grok: Three.js Site Build Showdown — Users compare local Qwen 3.8 27B, GPT-5.6 Terra, and Grok 4.6 on building a Three.js site. Source-reddit
  • OpenAI Launches ChatGPT for Teens with Safety Features — OpenAI launched ChatGPT for teens, adding new safety features for younger users. Source-reddit
  • Anthropic extends Claude Code weekly limits increase through August 31 — Anthropic extended the Claude Code weekly limits increase through August 31. Source-x
  • Theo defends T3 Code against inaccurate critical report — Theo pushed back against what he called an inaccurate critical report on T3 Code. Source-x
  • Deft AI Lab Launches Beta for Better Writing, Bypassing Detectors — Deft AI Lab launched a beta promising better writing that bypasses AI detectors. Source-x
  • GPT-5.6 Sol Gets 70% Discount in Devin Desktop, CLI — Devin is offering GPT-5.6 Sol with a 70% discount across the desktop app and CLI. Source-x
  • Nous Research Launches Bot Mode for Hermes Desktop with Customizable Bots — Nous Research launched Bot Mode in Hermes Desktop, enabling customizable bots. Source-x
  • Open-Source AI Job Search Tool Scores Listings and Tailors CVs — A new open-source tool scores job listings and tailors CVs automatically. Source-github
  • oMLX: LLM Server with SSD Caching and Continuous Batching on Mac — oMLX is an LLM server for Mac featuring SSD caching and continuous batching. Source-github
  • Chinese AI models approach quality of paid tools, prompting switches — Chinese AI models are approaching paid-tier quality, prompting users to switch providers. Source-reddit
  • pagedMark strips AI provenance from generated images and videos — pagedMark is a tool that removes AI provenance metadata from generated images and videos. Source-reddit
  • Harris Poll Study: Why People Love and Hate AI — A Harris Poll study unpacks the reasons people love and hate AI. Source-reddit
  • AQuA Paper’s Faulty Feature Fails Clean Re-Split — A clean re-split of the AQuA paper revealed the unusually strong result depended on a faulty feature. Source-reddit
  • Sainsbury’s Pauses AI Facial Recognition After Wrongful Shoplifting Accusation — Sainsbury’s paused its AI facial recognition system after a wrongful shoplifting accusation. Source-reddit
  • AI Coding Decline Was User’s Sloppy Prompts, Not Model Drift — An investigation finds declining AI coding quality was due to sloppy user prompts, not model drift. Source-reddit
  • Calls Grow for Mandatory AI Chatbot Disclosure — Critics are calling for companies to be required to disclose when users are interacting with AI chatbots. Source-reddit
  • Google Buys Bankrupt Spirit Airlines’ Data at Auction for AI — Google purchased bankrupt Spirit Airlines’ data at a bankruptcy auction for AI purposes. Source-reddit
  • Does AI automation actually save time or create more work? — A discussion asks whether AI automation actually saves time or ends up creating more work. Source-reddit
  • Claude Users Report Widespread Outage — Users report another widespread outage affecting Claude. Source-reddit

Generated by AI News Agent | 2026-08-18