TLDR AI 2026-09-11
Agents API π€, Cognition SWE-2 π¨βπ», Muse Shared Agents π§βπ€βπ§
Forge: The conference for companies making their own frontiers (Sponsor)
The most ambitious companies are training specialized models to outperform the frontier on their domain.
On November 3 in San Francisco, Fireworks Forge will bring together leaders making this shift.
At Forge, you'll
- Learn what it takes to stop renting and start owning your intelligence
- See how to take back control of your token bills with intelligent routing
- Join hands-on workshops for training, inference, routing, and evals
Hear from Jensen Huang, CEO of NVIDIA; Michele Catasta, President and Head of AI at Replit; Lin Qiao, CEO of Fireworks; and more speakers to be announced.
November 3 Β· Pier 27 Β· San Francisco
Attendance is free and limited.
π Apply to attend Forge
OpenAI Launches the Agents API (3 minute read)
OpenAI introduced the Agents API in public beta, giving developers access to the managed agent harness and infrastructure behind Codex. It handles context, tools, subagents, persistent execution, files, and code environments for agents that can run for extended periods.
Meta to announce Shared Agents for Muse at Meta Connect (2 minute read)
Meta's Muse app will introduce a "Shared Agents" feature, allowing users to create customizable agents that can be shared with others, similar to Grokbot's system. This could benefit small businesses already using Meta platforms by enabling specialized agents for workflows like customer support and sales.
OpenAI Is Open to Slowing Cutting-Edge AI, CEO Sam Altman Tells Staff (2 minute read)
OpenAI is considering slowing down development of cutting-edge AI. The company has raised concerns about its advanced AI systems, saying that model development should be paused due to safety concerns. The company's researchers have gone viral for making comments saying OpenAI's technology will kill humanity by the end of the decade.
π§
Deep Dives & Analysis
Does Scaling Web-Video Pre-training Help Real Robots Do Real Work? (36 minute read)
Larger video models and increased pre-training compute improve robot task performance, confirmed through Direct Video-Action models. Performance gains arise from better predictions of held-out web videos, with larger models excelling in real-world tasks like complex industrial unpacking. Pre-training quality, measured via DINO FD, predicts downstream robot efficiency, enhancing scalability for real deployments.
An operationalization of opaque serial depth (3 minute read)
Chain-of-thought (CoT) is a valuable tool for overseeing AI models. However, some architectural shifts could significantly reduce CoT monitorability. This study looks at how much unverbalized serial cognition a model can perform.
Detecting and countering misuse of AI: September 2026 (5 hour read)
Anthropic's Threat Intelligence team has identified and disrupted several operations in the past several months where threat actors tried to use Claude for malicious activity. This report shares case studies from those operations and describes how malicious use of Claude has evolved since the company's previous threat reports. The report covers disruptions between December 2025 and August. None of the misuse cases involved the use of Claude Fable or Mythos-class models.
π¨βπ»
Engineering & Research
π§ββοΈ Peace of mind in every sprint (Sponsor)
Writing code can be stressfulβbut not half as stressful as a surprise security meltdown. Inject optimism and calm into the developer scrum with
Microsoft Azure. Unified security across code and cloud environments and built-in DDoS protection mean you've got less cause for concernβand a clear mind for innovation.
Help secure your apps with Azure >Model Card for North Small Translate (8 minute read)
North Small Translate is an open-weights research release. It has 25 billion active parameters and 218 billion total parameters. The model is specialized for high-quality machine translation across 50 languages.
Introducing SWE-2: Pushing the Pareto Frontier (23 minute read)
SWE-2 pushes the Pareto frontier and achieves 50.0% on FrontierCode 1.1 Main1, while being 64% cheaper. It beats SWE-1.7 and Grok 4.6 on both score and cost, matches GPT-5.6 Sol and Fable 5/5.1 at a fraction of their price, and comes within a few points of GPT-6 Astra at a quarter of the cost.
OpenCodeReview (GitHub Repo)
Open Code Review is an AI-powered code review CLI tool. It originated as Alibaba Group's internal official AI code review assistant β over the past two years, it has served tens of thousands of developers and identified millions of code defects. It has been validated at massive scale. The agent can read full file contents, search the codebase, inspect other changed files for context, and produce deep reviews.
OpenAI launches GPT-Live-1 for full-duplex voice agents (2 minute read)
GPT-Live-1 is now available in the OpenAI API at $0.05 per minute. The model adds full-duplex speech, interruption handling, and 12 voice options. It can listen and speak at the same time, handle interruptions and acknowledgements as they happen, and keep conversations moving while performing deep reasoning or actions. The model can control tone, space, and style through the system prompt. Early tests report 80% fewer interruptions than with previous turn-based systems.
OpenAI puts Pro subscriptions on hold due to Astra demand (2 minute read)
OpenAI has paused subscriptions for its $200-per-month Pro plan. The company's Astra model is now rolling out to Pro, Plus, Enterprise, and Business accounts. The model promises a major leap forward in reasoning, coding, and computer use. OpenAI says the model is the beginning of the AGI era.
Why the world's best AI startups write bad prompts (& how to fix this) (20 minute read)
AI startups often accumulate sprawling prompts full of contradictions and ambiguity as teams continuously add instructions. Treating prompts like product and code, with modular sections for background, behavior, and output, can improve agent quality, reduce regressions, and lower costs.
AI agents that build lasting loyalty at Booking.com, SAP, and Microsoft (Sponsor)
Loyalty doesn't have to wait on hold. Parloa's AI CX agents instantly manage millions of conversations in any language. See how they build meaningful customer relationships.
Get a demoMeta's WearableQA Health Reasoning Benchmark (GitHub Repo)
Meta has released WearableQA, a benchmark with thousands of questions built from real-world wearable data, blood biomarkers, and demographics from 200 users.
Salesforce Finds Better Ways to Co-Evolve Agents and Their Harnesses (9 minute read)
Salesforce found that directly training smaller models on expert agent trajectories can hurt performance after their harness has already been optimized.
OpenAI Launches ChatGPT for Financial Services (4 minute read)
OpenAI introduced a financial-services version of ChatGPT Work, combining GPT-6 Astra with built-in premium data from popular providers.
Introducing the Google Cloud Developer Plugin for AI Coding Agents (4 minute read)
Google's new Google Cloud plugin is designed for AI coding agents, featuring installable bundles, agent plugins equip the AI agent of your choice with skills, and tools to be more effective on Google Cloud.
Universal Music is launching an AI music platform with ElevenLabs (2 minute read)
Universal Music Group is launching an AI platform with ElevenLabs, allowing users to create remixes and mashups using UMG's licensed music catalog.
Get the most interesting AI stories and breakthroughs delivered in a free daily email.
Join 1,100,000 readers for
one daily email