TLDR AI 2026-08-28
Gemini Omni 1.1 🎬, Cohere Parse 📄, Codex persistent mode 👨💻
One API for Every Model (Sponsor)
Access 500+ models from 80+ providers, with 200T+ tokens routed monthly through
OpenRouter.
> Best execution: intelligent routing (price / latency / quality) and Auto Router.
> Built for reliability: multi-provider failover and degradation routing.
> More control: per-request policies, team controls, observability and auditability.
> Easy to switch: OpenAI-compatible API; keep existing SDKs and switch models with a parameter. Start Building Now
Introducing Parse: Enterprise document intelligence at scale (5 minute read)
Cohere Parse converts complex, multimodal files into structured machine-readable data. A cost-effective vision language model for processing large volumes of enterprise documents, it detects and understands key visual elements and can process documents and images across nine major world languages. Parse can be accessed through the Cohere API for just $1.50 per 1,000 pages. A free version is available in Cohere Space for users to try.
Anthropic's Model Hardware Standard (4 minute read)
Anthropic opened a research preview of the Model Hardware Standard, a model-agnostic specification for letting AI agents operate scientific and manufacturing equipment.
Introducing H3 Max by fal (5 minute read)
H3 Max is a post-trained version of MiniMax H3 optimized for maximum speed. It can generate a 5-second-long video in under 3 seconds. The model is now available on fal for 50% off for the first week. This article discusses how the model was trained.
AI and the city (8 minute read)
AI is driving more distributed business formation in the US, with a rise in new companies forming outside major cities and in outer suburbs. This trend began during COVID-19 with remote work and accelerated with AI, yet business formation does not correlate directly with work-from-home levels. AI labs and product companies still cluster in major metro areas like San Francisco, highlighting the continued value of agglomeration for innovation.
Accelerating MiniMax-H3 (12 minute read)
A detailed benchmark of MiniMax-H3 video generation on 8× H200 GPUs showed SGLang Diffusion reaching 1.95x lossless speedups and as much as 6.24x with step reuse and sparse attention.
GPT 5.6 Discounts & Jevons Paradox (4 minute read)
OpenAI introduced large discounts on some of its models between July 27 and August 14. This resulted in Luna token usage jumping 13.8x and Terra token usage rising 5.6x while Sol, which remained at list price, only saw a bump of 1.1x. Most of the share gained by OpenAI's discounts came from other labs as opposed to cannibalization within the OpenAI family of models. Nearly a third of users who tried a discounted OpenAI model kept using it after the discounts expired.
Will the AI boom continue? Forecasting the trajectory of the AI industry (11 minute read)
Anthropic and OpenAI show impressive revenue growth, with expectations for continued rapid expansion. Forecasts indicate moderated growth for semiconductor stocks but an uptick for software stocks, with AI's impact on the software sector remaining a key consideration.
👨💻
Engineering & Research
What sets AI successes apart from the failures? (Sponsor)
Gartner has
the data. Every year, they interact with 10K+ CIOs and their teams, 80K+ business executives, and 5K+ technology providers. Their AI Hub already has over 4,000 use cases and case studies. You don't have to guess your way to AI success.
Explore the AI Hub.Putting Task Expertise into RL Achieves State-of-the-Art Performance on Text-to-SQL (18 minute read)
Adding scaffolding around text-to-SQL models improves performance somewhat, but this is limited by the base models' capabilities. This suggests that the task knowledge contained in the scaffold belongs in the training signal instead. Expert judgment needs to be part of the training process. Putting task expertise into task-specific training is ultimately what scales.
Fast On-Device Voice Cloning (9 minute read)
Halo Neuro introduced Sopro V2 and open-sourced Sopro V2 Turbo, a 120M-parameter multilingual voice-cloning model that can stream on laptop CPUs and in browsers.
Support persistent reasoning effort (2 minute read)
Codex has added 'persistent' to the reasoning-effort protocol and TypeScript SDK types. When a user targets a custom Responses-compatible provider whose model-defined effort is literally persistent, this now deserializes to the new Persistent variant and is unconditionally rewritten to disabled. Previously, it remained Custom("persistent") and was forwarded unchanged. Existing configurations and resumed sessions can therefore send a different value and change behavior or be rejected, so this alias should be limited to providers that define the translation.
Gemini Omni 1.1 Flash (4 minute read)
Google introduced Gemini Omni 1.1 Flash with new controls for extending scenes, interpolating first and last frames, 4K upscaling, and faster video iteration through the Gemini API.
Nvidia Climbing the Wall of Worries (7 minute read)
The average sell-side estimates for Nvidia's FY 2028 revenue a year ago were around $310 billion. Nvidia guided for around 70% revenue growth in FY 2028 yesterday, which translates to approaching around $700 billion in revenue next year. While analysts kept updating the estimates throughout last year, they were still short by around $125 billion. Nvidia's revenue growth of around 70% could be the floor for next year.
An update on AI's most important number (8 minute read)
OpenAI and Anthropic are seeing unprecedented growth, with revenues surpassing $100 billion combined, fueled by rapid adoption and technical advancements. This swift rise challenges expectations, as typical tech growth slows at such scales, potentially reshaping the economy if sustained.
Get the most interesting AI stories and breakthroughs delivered in a free daily email.
Join 1,100,000 readers for
one daily email