TLDR

TLDR Data 2026-09-03

Netflix’s Parallel Spark Tuning 📊, Tracing AI Token Waste 💸, S3 Gets More Duck 🦆

What do Uber, CERN, and Apple have in common? (Sponsor)

📱

Deep Dives

Running Apache Spark experiments in my sleep (and on a plane) (9 minute read)

How we eliminated $1 million a year of wasted AI agent spend in one hour (6 minute read)

Rerouting the Stream: How Lyft Moved to the Apache Flink Operator (11 minute read)

Read your own writes, off the primary (10 minute read)

🚀

Opinions & Advice

Why AWS Bought DuckLabs (8 minute read)

Building and Using Named Queries (7 minute read)

💻

Launches & Tools

MLCommons Releases New MLPerf Storage v3.0 Benchmark Results (4 minute read)

dbt doctor (GitHub Repo)

Introducing FrontierHarness Eval (14 minute read)

🎁

Miscellaneous

⚡️

Quick Links

Keenable SELECT showcase (Tool)

Curated deep dives, tools and trends in big data, data science and data engineering 📊

Join 590,000 readers for one daily email