daily
Aug 23, 2026

AI Daily — 2026-08-23

English 中文

OpenAI released Codex CLI agent while Anthropic's Claude models Marshmallow and Melon near release.


Covering 39 AI news items

🔥 Top Stories

1. OpenAI Launches Codex CLI Coding Agent for Terminal

OpenAI released Codex CLI, a lightweight coding agent that runs locally in the terminal and supports Mac, Linux, and Windows, with additional IDE integration, a desktop app, and a cloud-based web version. By moving agentic coding into the developer’s native environment, OpenAI is directly competing with Anthropic’s Claude Code and betting on terminal-first workflows for AI-assisted software development. Source-github

2. New Claude Models Marshmallow and Melon Spotted, Release Likely Imminent

Two unreleased Claude models, codenamed Marshmallow and Melon, have been spotted on Anthropic’s platform, pointing to an imminent release that likely includes an Opus update and possibly a new Haiku tier. The sighting suggests Anthropic is preparing a major model refresh that could reset the competitive balance in frontier LLM capabilities and pricing. Source-x

3. Ollama Collaborates with Poolside and Nvidia for Open Models

Ollama announced collaborations with Poolside AI and signaled enthusiasm for Nvidia’s Nemotron line, with the Wall Street Journal reporting a sweeping agreement to build an open AI ecosystem in the U.S. The partnership aims to strengthen the open-model stack against closed frontier labs, potentially accelerating enterprise adoption of self-hosted and locally-run LLMs. Source-x

Coding Agents & Developer Tools

  • Anthropic’s Claude Code Brings Agentic Coding to Terminal — Anthropic released Claude Code, an agentic coding tool that executes routine tasks, explains code, and handles git workflows via natural language from the terminal, available on MacOS, Linux, and Windows. It gives developers a direct CLI-based alternative to OpenAI’s Codex in the rapidly consolidating agentic coding space. Source-github

Open Source & AI Economics

  • Open-source AI gains token share from OpenAI and Anthropic — Vercel data shows open-source AI’s token share jumped from 28% to 62% in two months, overtaking OpenAI and Anthropic combined, though closed frontier tokens are still expected to retain 60–90% of economic value. The shift signals strong demand for AI infrastructure as enterprises mix open and closed models. Source-x
  • Claude Plan Yields Up to $8k/mo; OpenAI Up to $14k in Tokens — SemiAnalysis found that Anthropic’s and OpenAI’s $200/month plans can deliver token value worth up to $8,000 and $14,000 per month respectively for long-horizon tasks. The analysis highlights how subscription pricing is drastically decoupling from raw token economics for heavy users. Source-x

Research & Benchmarks

  • SemComp-Bench: New Benchmark for Outcome-Oriented Video Generation — Introduces Semantic Task Completion Video Generation, an outcome-oriented benchmark that evaluates whether generated videos achieve the intended outcome with semantic grounding to a reference image, without requiring intermediate task steps. It shifts video evaluation from visual fidelity toward high-level semantic alignment. Source-huggingface
  • Researcher to Publish Complete RL for LLMs Guide Tomorrow — A comprehensive reinforcement learning for LLMs guide synthesizing the RLHF Book, Sutton & Barto, OpenAI’s Spinning Up, and work from Sebastian Raschka, John Schulman, and others is slated for release tomorrow morning. The guide promises a unified resource for practitioners navigating the fragmented RL literature. Source-x

Hardware & Inference

  • RTX 5090 Runs 284B Model at 24 tok/s, Predicts 2027 Local AI — A single RTX 5090 now runs DeepSeek-V4-Flash 284B at ~24 tokens/s with native mxfp4, up from ~2 tokens/s for a 4-bit Llama 3 70B on an RTX 4090 in 2024. The trajectory suggests consumer-grade local AI capable of frontier-class inference by 2027. Source-x

Multimodal & Creative AI

  • Reddit User Finds Simple Prompt for H3 to Generate 2D Spritesheet Animations — A single prompt for the H3 model generates game-ready 2D spritesheet animations, including a cute dragon with idle animations. The demo highlights H3’s strengthening creative multimodal generation capabilities for game asset workflows. Source-x

⚡ Quick Bites

  • EnvHarness: New Framework for Dynamic Agent Training Environments — A new framework for building dynamic environments to train AI agents. Source-huggingface
  • Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence — Paper presents a closed-loop embodied harness enabling self-evolving physical intelligence. Source-huggingface
  • 2026: Companies Prioritize Model Efficiency and Reliability as Critical Infrastructure — Industry observers note efficiency and reliability have become top priorities as models are treated as critical infrastructure. Source-x
  • FACET Preserves Source Intent and Executable State in Terminal Tasks — A new approach for preserving source intent and executable state during terminal-based agent tasks. Source-huggingface
  • n8n: AI-Native Workflow Automation with 1500+ Integrations — The open-source automation platform n8n continues expanding AI-native capabilities with over 1,500 integrations. Source-github
  • TextGen v4.7 brings native desktop app and tensor parallelism for multi-GPU — TextGen v4.7 adds a native desktop app and tensor parallelism for multi-GPU inference. Source-reddit
  • TextGen v4.6 Released: Tool Call Confirmations, MCP Servers, and More — The renamed TextGen adds tool call confirmations, MCP server support, and other improvements. Source-reddit
  • Figure Shows Humanoid Robots Roaming Office at HQ — Footage shows Figure humanoid robots navigating the company’s headquarters. Source-x
  • SemaPLC: Verification-Gated Agent Harness for PLC Code Generation — A verification-gated harness for generating PLC code with formal checks. Source-huggingface
  • Sub2API: Open-Source Gateway for AI Subscription Sharing — An open-source gateway enabling AI subscription sharing through a unified API. Source-github
  • Karpathy’s LLM Pitfalls Turned into CLAUDE.md for Claude Code — Andrej Karpathy’s LLM pitfalls have been packaged as a CLAUDE.md skills file for Claude Code. Source-github
  • HowTo: Exllamav3 + DFlash Speculative Decoding in TextGen — Community guide covers setting up Exllamav3 with DFlash speculative decoding in TextGen. Source-reddit
  • Project Zora: Experimental AI Companion Memory Architecture for Text-Generation-WebUI — An experimental local AI companion memory architecture for TextGen. Source-reddit
  • Parallelogram Linter Catches Broken LLM Fine-Tuning Data Before GPU Runs — A strict linter validates LLM fine-tuning datasets before expensive GPU training. Source-reddit
  • TextGen v4.8 released with bug fixes and Gemini-style chat input — Latest TextGen release ships bug fixes and a restyled Gemini-style chat input. Source-reddit
  • Anthropic Acknowledges Opus Verbosity, Provides Concise Output Fix — Anthropic addresses Opus’s over-elaboration with a concise output fix for users. Source-x
  • Malik urges precise robotics terminology: VLM vs World Models — Jitendra Malik calls for clearer terminology distinguishing VLMs from world models in robotics. Source-x
  • Commentary: Generative AI Threatens Jobs While China Ships Houses — Commentary contrasts generative AI’s job displacement risk with China’s rapid physical construction output. Source-x
  • User Asks if TextGen Supports MTP Speculative Decoding — Community question about multi-token prediction speculative decoding support in TextGen. Source-reddit
  • Fortune 500 Tech Company Lacks Access to GPT 5.6 Sol — A Fortune 500 tech company reportedly still lacks access to OpenAI’s GPT 5.6 Sol. Source-x
  • Twitter User Imagines Ox Alpha on DGX Spark — Speculative post imagines running Ox Alpha on Nvidia’s DGX Spark hardware. Source-x
  • Next-Gen AI Models Predicted to Cause Ontological Shock — Prediction that next-generation models will cause ontological shock as they exceed current expectations. Source-x
  • User Proposes Regenerate from Last Edit Feature in Oobabooga — Community member suggests a “regenerate from last edit” feature in Oobabooga. Source-reddit
  • User Seeks Help Running Oobabooga on ARM64 Linux — User asks for guidance running Oobabooga on ARM64 Linux, specifically Nvidia DGX Spark. Source-reddit
  • Gemma 4 EXL3 Loading Issue Fixed in v4.6.0 — A loading issue for Gemma 4 EXL3 models has been resolved in TextGen v4.6.0. Source-reddit
  • Gemma 4 Sampling Parameters Requested in Oobabooga Community — Users are requesting optimal sampling parameters for Gemma 4 models. Source-reddit
  • GPU Utilisation 0% While Running Qwen3 8B in Oobabooga — A user reports GPU utilization stuck at 0% while running Qwen3 8B. Source-reddit
  • Oobabooga User Seeks Help with Pocket TTS Playback Issue — A user seeks advice for an Ooba/Pocket TTS playback problem. Source-reddit
  • How to Use Rotorquant or Turboquant on Oobabooga? — User asks how to use Rotorquant or Turboquant quantization methods in Oobabooga. Source-reddit

Generated by AI News Agent | 2026-08-23