TLDR AI 2026-09-07
Claude proves Fermat 🧮, automated AI researcher 🔬, Z1 efficiency chip ⚡
OpenAI Researcher Warned About Rapidly Advancing AI (17 minute read)
An OpenAI researcher says that reasoning models could continue advancing rapidly enough to contribute to their own development, creating increasingly serious alignment and cybersecurity risks.
Formalizing Fermat's Last Theorem (11 minute read)
Claude successfully created the first complete computer-verified proof of Fermat's Last Theorem in 11 days using Lean, automating the complex task initially proven manually by Andrew Wiles in 1995. The proof, verified via Prove2Me and Lean, involved 13 million lines of code and proved 29,500 intermediate theorems, proving AI's potential to ease the traditionally laborious formal verification of mathematical proofs.
Grok Imagine Video 1.5 agent (1 minute read)
Grok Imagine Video 1.5 agent delivers higher quality, better storytelling than previous releases. Powered by Grok's latest Image 2.0 model, it excels at connecting multiple shots together with greater continuity. The agent is now live on the web, iOS, and Android. A short video generated by the model is available in the thread.
AI Safety Is Not the Same as Security (6 minute read)
Frontier labs may be applying probabilistic AI safety techniques to problems that require deterministic security controls. Recent agent sandbox escapes highlighted the gap between reducing harmful model behavior and reliably containing software.
GPT‑6 Astra on robotic manipulation (15 minute read)
Researchers gave GPT-6 Astra control of YAM arms under an Inspect Robots agent policy and gave it two tasks: it had to pick up a red block from a table and place it inside a bowl, and pick up a round blue puzzle piece by the knob at its center and place it into the matching circular groove in the board. Astra placed the block in 19 of 20 trials in the bowl task. It completed the puzzle insertion two times in 20. Astra completes the bowl task far more often than Fable at about half the cost per run.
👨💻
Engineering & Research
Loyalty shouldn't wait on hold (Sponsor)
Your customers don't want to waste time. Parloa's AI agents
eliminate holds by managing millions of conversations, so your support team can focus on higher-value relationship building. Take for example, Decathlon, they eliminated 20% of their repetitive tasks with Parloa. Learn how,
request a demoLLM-as-a-Verifier (GitHub Repo)
LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and medical agentic benchmarks. The feedback generated from the framework can be used for test-time scaling, progress tracking, and reinforcement learning.
Random Attention (GitHub Repo)
Random Attention kept a uniformly sampled subset of generated KV-cache entries instead of relying on learned importance signals or attention statistics. Across several reasoning benchmarks and model families, it matched or exceeded more complex eviction methods while reducing eviction overhead.
Research acceleration: The view inside OpenAI (9 minute read)
OpenAI plans to develop an automated AI researcher by March 2028, aiming to enhance research efficiency while maintaining human oversight to ensure alignment and safety. Researchers now use coding agents more frequently, with increased code generation and experiment execution, shifting focus to more complex tasks. The organization paused reinforcement learning training temporarily following a security breach but continues to adapt safety measures and transparency to uphold the development of safe AGI.
OpenAI's AGI number came from a harness, not the model (6 minute read)
OpenAI claimed it had achieved AGI due to its 99.9% score on ARC-AGI-3. However, tests that ran the same model through the benchmark's own software scored 62.7%. The gap comes from the software around the model that OpenAI built. The different scaffolding around its agents helped OpenAI achieve the high score.
Anthropic IPO launch shifts toward mid-October (3 minute read)
Anthropic is expected to complete its IPO listing days before the US midterm elections in November. It will begin marketing the offering in mid-October at the earliest. The IPO prospectus will likely be released in late September. The plans could change - such changes are not unusual, as companies frequently have to adjust schedules due to market conditions, regulatory reviews, and other preparations.
Product Manager, Applied AI at TLDR ($200k base + $60k bonus, Fully Remote)
TLDR is hiring its first PM to help build the agent-first operating layer used across the company. We're looking for a builder who has shipped real products/systems with LLMs.
Click here to learn more.
An Interview with OpenAI President Greg Brockman About Astra and Alignment (63 minute read)
OpenAI President Greg Brockman discussed Astra, OpenAI's new model, focusing on its enhanced capabilities and alignment. He highlighted the importance of scaling infrastructure and addressed challenges in cybersecurity following the Hugging Face incident. Brockman shared insights on OpenAI's positioning within the tech value chain, emphasizing strategic focus on sectors like health and collaboration with partners like Nvidia and Microsoft.
Concrete, Silicon, & Leverage (4 minute read)
The US data center capacity will expand from 25 to 70 gigawatts, requiring $5 trillion, mostly financed by debt. This expansion creates a 34% growth in the US corporate bond market and raises questions about financing, potentially involving municipal bonds. To service this debt, annual AI revenue must grow from $150 billion to at least $1.2 trillion by 2030, requiring a 55% annual growth rate.
Get the most interesting AI stories and breakthroughs delivered in a free daily email.
Join 1,100,000 readers for
one daily email