TLDR AI 2026-08-03
DeepSeek V4 Flash ⚡, OpenAI’s math breakthrough 🔢, Qwen 3.8-Max 🤖
WorkOS MCP: Manage your auth platform from any AI agent (Sponsor)
Debugging SSO, managing users, adjusting auth policies, configuring branding: every configuration task lives behind a UI that only a human can drive. Until now.The
WorkOS MCP server gives agents the same access as your dashboard login.
✅ Hundreds of operations, discoverable at runtime.
✅ Connect in one command via OAuth, with scoped tokens instead of a master API key.
✅ Pass a screenshot of your marketing site and ask your agent to match the login page.
If a human had to do it before, an agent can do it now. Connect your agent →
Qwen3.8-Max: A New Bar for Coding and Cowork (25 minute read)
Qwen 3.8-Max is now available. The open weights will be released next week. The model, which has 2.4 trillion parameters, delivers comprehensive improvements across coding, work, research, and long-horizon tasks. It can answer questions as well as complete complex tasks end-to-end with greater reliability.
Ten advances in mathematics and theoretical computer science (4 minute read)
OpenAI has shared a selection of ten results discovered while evaluating an unreleased model. Each resolves or makes substantial progress on a long-standing open problem. These problems span high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography and extremal combinatorics. All of these problems are of substantial interest to their respective mathematical communities. Several are of broad interest across mathematics as a whole.
)
Microsoft tests new MAI Realtime voice model (2 minute read)
Microsoft's first native real-time voice model, MAI Realtime, has surfaced as a hidden early-access entry in the company's MAI Playground. The listing points to a bidirectional, full-duplex system that can listen and speak at the same time rather than trading turns. There are two voices available, both noticeably more natural than what Copilot's voice mode currently delivers. The model will likely be made available on Microsoft Foundry and Copilot voice, but no timeline is available.
👨💻
Engineering & Research
smevals (GitHub Repo)
smevals is a framework for running evals against AI models, small and large. An Eval is a collection of Tasks used to determine how good a particular model or model-and-harness configuration is at a specific high-level capability. Tasks are the individual exercises that a model must complete for its abilities to be evaluated. Evals can optionally be grouped into Suites of related Evals, primarily as a mechanism for organizing them on disk.
Ramp SWE-Bench (3 minute read)
Ramp built a private benchmark from 80 production backend tasks spanning payments, accounting, procurement, treasury, and fraud. It scores review-ready patches that pass tests within 45 minutes, exposing model trade-offs across accuracy, latency, and cost without public-benchmark contamination.
MSLK kernel reference (Website)
MSLK (Meta Superintelligence Labs Kernels) is a library of fused GPU kernels for transformer workloads. It contains a collection of high-performance kernels and optimizations built on top of PyTorch primitives for GenAI training and inference. MSLK is released in accordance with the PyTorch release schedule. There is no guarantee that each release works in conjunction with PyTorch releases that are older than the one that the MSLK release corresponds to.
TLDR is hiring a curator for TLDR AI! (TLDR Curator, ~5 hrs/week)
Over 1M subscribers read TLDR AI to stay on top of the latest in AI models, research, engineering, and more. If you work in AI and want to help curate it, send your LinkedIn or resume to
ai@tldr.tech!The Math Superstar Who's Terrified of AI—and Just Took a Job at OpenAI (9 minute read)
Jacob Tsimerman, who recently won the Fields Medal, is starting a position at OpenAI. Tsimerman previously wrote a paper categorizing the ways AI might kill everyone. He appears to be so worried about the dangers of AI that he's pivoting to work on AI safety. The star professor wants to use math to advance the study of AI and ensure that the technology won't lead to our extinction.
A new era of AI testing (2 minute read)
Andrej Karpathy asked Opus 5 to make a Three.js render of the first paragraph of The Lord of the Rings with a 1-million-token budget, and the model returned 5,500 lines of code that procedurally rendered the story. The LLM orchestrated the polygon assets and wrote code that animated it all according to the story. No human would have the stamina and patience to write something this custom, so they are a good test of what LLMs are capable of. A video of the generated animation is available in the post.
Get the most interesting AI stories and breakthroughs delivered in a free daily email.
Join 1,100,000 readers for
one daily email