TLDR AI 2026-09-09
Images 2.5 🎨, OpenAI solves Navier-Stokes 🧮, AlphaGenome Atlas 🧬
Your blueprint for AI governance (Sponsor)
Many organizations have AI policies, but far fewer have a practical way to decide whether an agent is ready for production, who signs off, and how that decision gets documented.
Get the new AI governance ebook to see how leading teams are building structured, evidence-based review gates for AI agents to connect evaluations, traces, human review, and approval workflows into an auditable record of deployment readiness.
Download the guide to learn how to:
- Centralize traces, model calls, evals, and approvals in one auditable record
- Evaluate AI agents in high-impact domains using structured assessments and expert review
- Build a governance program informed by ISO 42001, the EU AI Act, and the NIST AI RMF
Get the guide.
An OpenAI Model Solved the Navier–Stokes Millennium Problem (5 minute read)
OpenAI announced that an internal AI system produced a proof resolving the roughly 90-year-old Navier–Stokes existence and smoothness problem, one of mathematics' seven Millennium Prize Problems. The model showed that smooth three-dimensional fluid dynamics can develop a finite-time singularity and produced both an analytical proof and a Lean formalization.
Introducing Muse: The World's First Personal AI Agent Built for Everyone (5 minute read)
Meta introduces Muse, a personal AI agent, powered by Muse Spark, to help users achieve goals by automating tasks like booking travel or sending emails. Muse operates securely on Muse Secure VM, ensuring data privacy with unique protections like the Sentinel agent overseeing actions. Muse will soon offer encrypted data with Muse Confidential VM and is available on iOS, Android, and muse.ai in the US.
ChatGPT Images 2.5 (9 minute read)
OpenAI introduced ChatGPT Images 2.5 with sharper details, better reference-image preservation, more reliable editing, and up to 50% lower generation latency.
>10x More Efficient Pretraining (15 minute read)
Without large amounts of compute, small labs can only compete through algorithmic efficiency. Magic's pretraining recipe is now more than 10 times more compute-efficient than that of leading open-weight base models. The startup believes that pretraining, agentic RL, and long-context are sufficient for building superhuman coding agents and automating AI research and development. This post discusses its pretraining and long-context work.
Inside the megakernel serving engine for North Mini Code (22 minute read)
This post presents a fully fledged serving system built around a decode megakernel. The system supports everything a real server needs: continuous batching, paged attention, and ragged sequence lengths, all behind an OpenAI-compatible endpoint with tool calling. The megakernel reaches 292 tokens per second on batch size 1, or 62% of SoL - 1.58× faster than vLLM. That margin holds across batch sizes and out to 256K of context with no measurable loss of accuracy.
Pretraining progress is mostly coming from data (17 minute read)
Between 2019 and 2025, 3.24x more compute efficiency gains have come from data improvements rather than model improvements. The gains from data and model improvements are mostly independent and don't interact. Most model research has consisted of removing or pushing back constraints to scaling. The data improvements may matter less for larger models. Small models see significant gains from data quality.
👨💻
Engineering & Research
GPUs delivered with the inference stack flashed. AMD Instinct™ Coder. (Sponsor)
Building an inference stack takes months, especially in today's severely supply constrained environment.
AMD Instinct™ Coder, powered by Spectro Cloud, ships as a Supermicro server with AMD Instinct GPUs and
PaletteAI Inference Launchpad already flashed. Serve inference in a day.
Google's AlphaGenome Maps 9 Billion Genetic Variants (4 minute read)
Google DeepMind introduced AlphaGenome Atlas, a 1-petabyte database predicting the regulatory effects of all 9 billion possible single-nucleotide variants in the human genome.
Hyper-𝜏-bench: Evaluating agents that build agents (4 minute read)
Hyper-𝜏-bench places a developer agent into a sandboxed workspace with the records of a simulated business and a simulated client that it can message at any time. The developer agent recovers the spec from the evidence, designs the architecture, and turns the business' actions into tools until it has a working customer-service agent. The finished agent has to serve from a fixed menu of models within a cost budget per conversation. Claude Opus 5 (max reasoning) running in Claude Code passes just 23.9% of the held-out evaluation tasks when working alone. Paired with an engineer with deep context, the same class of model reaches 82.2% on the same tasks.
Introducing Mercury 2.5 (5 minute read)
Mercury 2.5 is the largest diffusion language model ever trained. It performs comparably to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. The model outputs 1,107 tokens per second on widely available Nvidia GPUs and has a 260K-token context window. At launch, Mercury 2.5 is 80% off at $0.04 per million input and $0.15 per million output.
Is the 3x AI Productivity Gain just a Computer that Never Sleeps? (3 minute read)
OpenAI researchers now supervise 3.14 agent-workdays per eight-hour shift, suggesting AI productivity increasingly comes from parallel, around-the-clock machine labor rather than less human effort. That leverage is expensive, with median daily inference spend rising from $14 to over $600.
I Asked 100 Agents to Hack Me (9 minute read)
Around 100 self-hosted agents attempted to hack various online accounts over five hours. They compromised three accounts through software vulnerabilities and two through password brute-forcing, while also making 16 social engineering attempts. The experiment tested abliterated open-source models' capabilities, revealing potential future risks as these models improve and become cheaper to deploy.
Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market (2 minute read)
Cognition has raised $2 billion at a $48 billion valuation in a funding round led by Andreessen Horowitz, Accel, Founders Fund, General Catalyst, and Avenir. The startup's soaring valuation signals that VCs still see room for multiple major players to capture meaningful market share in AI coding. Cognition leases an Nvidia server cluster that could push its total cash burn to $800 million this year. The startup is expected to reach $4 billion to $5 billion in annualized revenue by the end of 2026.
Get the most interesting AI stories and breakthroughs delivered in a free daily email.
Join 1,100,000 readers for
one daily email