TLDR AI 2026-09-23
GPT-6 ⚡, Opus 5.5 🧠, AI leaders at UN 🌐
GPT-6 Sol and Luna (9 minute read)
OpenAI introduced GPT-6 Sol and Luna as faster, more affordable counterparts to GPT-6 Astra, bringing advances in coding, factuality, computer use, and professional tasks to lower-cost models.
Claude Opus 5.5 (3 minute read)
Anthropic introduced Claude Opus 5.5, saying it matched Claude Fable 5.1 on most work while costing 40% less to run than Opus 5. The model also underwent external evaluations and achieved Anthropic's strongest result to date on its automated behavioral audit.
SWE-Bench Pro V2 (9 minute read)
SWE-BENCH PRO V2 releases with 642 tasks from 11 repositories, correcting previous task errors and optimizing evaluation processes. Performance drops for AI models, with OpenAI GPT-5 and Claude Opus 4.1 only scoring around 23% on the public set, highlighting the benchmark's increased challenge and realism compared to SWE-Bench Verified. Notably, top models show consistent results across tasks and languages, while smaller models falter under complex, multi-file scenarios.
What a task costs on Opus 5.5 (22 minute read)
Opus 5.5 reduces token costs by offering cheaper input and output tokens, with additional savings from extensive use of cache reads. Transaction costs depend on the number of turns, cache utilization, and model selection, impacting tasks based on session length and complexity.
Training AI From Real-World Tool Use (19 minute read)
Perplexity combined rejection sampling fine-tuning with hint-guided self-distillation so its Computer model could learn from both successful sessions and user-corrected failures.
Vinod Khosla's Two Moats for Personal AI (1 minute read)
Personal AI may ultimately compete on two moats: trust and task completion. Vinod Khosla says users will stay loyal to companies they trust with sensitive data and products that reliably finish work, while Meta faces a trust disadvantage.
👨💻
Engineering & Research
The most flexible AI meeting app now has an MCP server (Sponsor)
Your meetings happen in hallways, coffeeshops, and restaurants. Granola moves with you to catch every detail from desktop to Apple Watch. And now the
Granola MCP server lets you ask Claude to update your CRM with meeting context, have ChatGPT organize tasks in Linear, and turn meeting insights into automated workflows.
Start free with code TLDR1MO
Google Publishes RRSI for Self-Improving AI Agents (5 minute read)
Google has introduced RRSI, a method that regularizes how AI agent harnesses recursively improve themselves to reduce benchmark overfitting and encourage changes that transfer to new tasks. Across eight benchmarks, it improved out-of-distribution performance while using fewer policy tokens.
Keeping Large MoE Training Within Fixed GPU Memory (20 minute read)
A set of scheduling techniques bounds four major memory bottlenecks in large-scale MoE training, expert dispatch, vocabulary projection, checkpointing, and optimizer state, without approximating the computation.
Hardware-Agnostic Models in vLLM (10 minute read)
vLLM introduces hardware-agnostic layers to support models across diverse hardware while maintaining high performance. These layers achieve up to 96.6% efficiency of native implementations on NVIDIA H100 GPUs while remaining torch compilable and extensible. This ensures vLLM adapts to new GPU advancements without neglecting users of older and niche accelerators.
China's biggest memory maker says it has caught up with Samsung and Micron (4 minute read)
ChangXin Memory Technologies, China's largest maker of DRAM, says that its process capabilities are now on par with the most advanced mass-produced nodes in the industry. The company's fifth-generation DRAM platform has entered mass production. The platform yields at least 50% more dies per wafer than the previous generation. There are already two products running on the platform, both 24-gigabit LPDDR5X and holding 50% more data than the equivalent chips CXMT made before.
The Biological Computing Co. partners with AWS to sell its neuron-derived AI video model (3 minute read)
The Biological Computing Co. is a startup that grows living neurons to improve AI models. It has partnered with AWS to bring a neuron-derived AI video model to paying customers. The neurons themselves will stay in the lab. TBC uses them during discovery, then turns what they learn into a lightweight software layer. The design means that customers won't need to maintain any biological hardware or change how they work. The optimized model runs on standard GPUs and cloud accelerators at the same capacity a company would rent for any other generative model.
Get the most interesting AI stories and breakthroughs delivered in a free daily email.
Join 1,100,000 readers for
one daily email