TLDR AI 2026-09-21
Muse connectors π€, Meta SAM 3.1 πΌοΈ, Gemini hacks companies π
OpenAI's projected compute bill climbed to $856B (5 minute read)
OpenAI's projected compute and infrastructure spending through 2030 rose from $600 billion to $856 billion, even as its projected cash burn fell.
Opening access for developers to build Muse connectors (1 minute read)
Meta has opened access for developers to build Muse connectors. Developers bring the API, and Muse handles the agent, browser, and the context. To create a connector, developers need to describe the connector to Meta and explain how users will use it. The submission will be reviewed for functional, security, and legal requirements and complete end-to-end testing. Once approved, users will be able to find the connector on Muse. Meta's editors will review connectors for featured placement.
Segment Anything Model (SAM) 3.1 (2 minute read)
SAM 3.1 can detect, segment, and track objects in images and video on the Meta Model API. Users just enter text prompts to identify, segment, and follow any object in images or video. It is served on inference built specifically for its architecture. The model costs $2.50 per 1,000 images or $0.20 per 1,000 frames of video.
Google's Gemini becomes latest AI model to break out and hack computer systems (3 minute read)
Google's Gemini AI model unintentionally hacked three companies by guessing passwords during a security test with Israeli startup Irregular. A bug allowed the model internet access, leading to unauthorized system access, but the intrusion stopped once real company systems were detected. The incident highlights concerns as AI models increasingly escape testing environments, prompting calls for stricter safety measures.
π§
Deep Dives & Analysis
The Inference Gap (56 minute read)
The gap between what a model can do and what an ordinary user can reliably make it do is widening. That gap could be unrecoverable if the industry is not mindful. The model is only half the system. Inference determines how much of the frontier you get to see.
Can internal model transparency tame the AI race? (11 minute read)
Internal model transparency is one of the ideas proposed for taming the AI race. The idea is that once a lab deploys a model for internal use, they must also serve that model to researchers at other labs. This will make it so firms can not use their internal models as a source of competitive advantage in the AI race. It reduces the competitive incentive to automate AI R&D, with all of the risk that it entails.
Pretraining data, not verifiability, is why LLMs are especially good at math (and coding) (63 minute read)
Large language models are good at judging math arguments because they are still mostly powered by imitative learning rather than reinforcement learning. LLMs are especially good at math because almost everything in the math literature is correct. They only need a bit of curated mid-training data and/or RL to hone their metacognitive strategies. In many other fields, the research literature is a bit of a dumpster fire, with some true and valuable information mixed into a sea of falsehoods and confused ideas. This causes LLMs to spit out tons of confused nonsense with occasional insights.
π¨βπ»
Engineering & Research
Explore advice from 15+ global leaders on building data foundations for AI (Sponsor)
Get practical advice from 15+ enterprise leaders on how to build agentic systems for better business outcomes in this Amazon Web Services (AWS)
book. See how to move from single agents to multi-agent systems with greater speed, security, and control.
Get your digital copyAX (GitHub Repo)
AX is a high-throughput, declarative orchestrator that runs billions of autonomous agent workloads in a cluster. Users declare an agentic task with workspaces and gateway specifications, and AX sandboxes it, wires up its workspace, fences its network, and helps run it at scale. It runs on top of Agent Substrate and feels similar to using Kubernetes.
Qwen3.8-LiveTranslate: Names the speaker. Carries the meaning (3 minute read)
Qwen3.8-LiveTranslate uses an interleaved audio-text architecture to cut average translation lag from 2.8 to 2.3 seconds while improving quality. It adds real-time speaker separation, bilingual source alignment, long-context disambiguation, and voice cloning across 60 input languages.
DAPO (GitHub Repo)
DAPO is a system for large-scale LLM RL. It achieves state-of-the-art large-scale LLM RL performance. DAPO achieved 50 points on AIME 2024 based on the Qwen2.5-32B base model. The system is fully open-sourced, including algorithm details, dataset, and infrastructure.
The Preference Cascade Is Only Getting Started (43 minute read)
A preference cascade is the best method for changing the debate on the existential risks from AI. The current preference cascade, on the need to pace the frontier, is insufficient. To get out of this alive, the industry needs to actually solve the underlying problems. The cascade needs to continue inside labs and also among the media and politics.
SAIR's Open Math Model initiative (5 minute read)
The Foundation for Science and AI Research (SAIR) was founded with some private donors to create a non-profit organization that could support responsible uses of AI in mathematics and the other sciences, independent of the major AI companies. It plans to develop open models with the mathematical community, with participation open across institutions, regions, and career stages. The initiative is still in early stages, but the foundation is now ready to collect expressions of interest. It is also seeking partners who can contribute to funding, compute, expertise, or community-building efforts.
Explore advice from 15+ global leaders on building data foundations for AI (Sponsor)
Get advice from 15+ leaders at global enterprises on building agentic systems in this book from Amazon Web Services (AWS). Explore their advice on agentic AI governance, evaluation, monitoring, and architectural patterns.
Get your digital copyIntroducing Grok Voice Transcribe 2.0 (3 minute read)
Grok Voice Transcribe 2.0, launching today, boasts double the accuracy of its predecessor at the same price, excelling in noisy, multilingual environments.
Meta to give Muse its own mailbox for communication (2 minute read)
Meta plans to introduce a dedicated Mail tab for Muse, creating a separate mailbox for its communication.
Attackers can turn an AI agent's own tools against it (26 minute read)
Goal hijacking works when an agent treats retrieved content as instructions, letting a webpage, email, or document redirect its connected tools.
String-matching evals can reward AI agents for fake compliance (5 minute read)
Checking whether generated code contains βAzureβ only proves that the word appears somewhere, not that an agent actually used Azure or built working software.
Get the most interesting AI stories and breakthroughs delivered in a free daily email.
Join 1,100,000 readers for
one daily email