TLDR Data 2026-10-05
Cloudflare’s Decision Models 🎼, Inside OpenAI’s Data Agent 🧠, DuckDB Inside Aurora 🔗
Diving Through Data at OpenAI: How a Data Agent Navigates 70,000 Datasets (27 minute video)
OpenAI built an internal AI data agent that answers complex data questions by combining warehouse queries with rich context from schemas, lineage, code, company knowledge, memory, and live sources, while showing its assumptions, SQL, citations, and confidence. The big lesson is that reliable data agents depend less on raw model intelligence and more on giving them the right context, continuously learning from corrections, and testing them with strong evals.
How we extended Apache DataFusion to execute one query across many machines (13 minute read)
Datadog built Distributed DataFusion to let a single Apache DataFusion query run across multiple machines, giving heavy interactive queries much more compute while keeping lightweight queries on one machine to avoid unnecessary overhead. In production, some expensive queries ran up to 10× faster at the same total cost. Datadog released the system as an open-source framework that teams can adapt to their own data sources and infrastructure.
How LinkedIn extended ClickHouse from distributed tracing to metric discovery and analytics (12 minute read)
LinkedIn replaced three systems for finding and analyzing metric metadata with a single ClickHouse index. It now handles more than 150,000 queries a minute across 13 billion metrics, averaging 68 milliseconds per query. The team explains how it migrated behind an existing gateway and cut memory use to about a fifth of the old setup.
Context Is the New Code (50 minute video)
Context is becoming a primary software artifact for AI coding agents: teams must generate, lint, version, and test context packages with the same rigor as application code. Managing context drift, establishing AI bills of materials, and instrumenting production retrospectives are now essential data engineering disciplines to prevent agent hallucinations and silent architectural degradation.
LinkedIn Built Kafka. What Did It Change When Kafka Was No Longer Enough? (9 minute read)
LinkedIn outgrew Kafka's partition model at extreme scale and built Northguard to separate ordering, replication, and placement, making data movement and recovery less painful. Xinfra then hid the migration behind a compatibility layer so applications could move gradually, showing how infrastructure rewrites become practical when clients are insulated from backend changes.
Order from Chaos: Strategies for unifying structured and unstructured data (Sponsor)
Learn how to analyze structured and unstructured data (like text, images, and PDFs) together in BigQuery, without moving data. Incorporate techniques like multimodal processing, automatic embedding generation, and hybrid search; then go further with generative AI functions and ObjectTables to extract deeper, unified insights.
Watch the webinar session or
read the blog.
Amazon Aurora PostgreSQL now supports direct querying of Apache Iceberg and Parquet data in your data lake (5 minute read)
Amazon Aurora PostgreSQL now queries Apache Iceberg and Parquet tables directly in Amazon S3 data lakes without requiring separate ETL or reverse-ETL pipelines. Powered by an embedded DuckDB engine, the feature supports standard PostgreSQL syntax, predicate pushdown, and column pruning across operational and historical analytical data with zero additional Aurora licensing fees.
Introducing Clef: our open-source decision models, and new RL fine-tuning platform (10 minute read)
Cloudflare has released Clef and Clef-flash, ‘Jev-like' open-source System 1 models that turn text or images into choices and probabilities. They can help route support tickets, classify websites, or decide when an agent should ask a human. Both run on Workers AI or locally, and Cloudflare is offering help fine-tuning them for specific tasks.
Apache Iceberg 1.12.0: What's New, Breaking Changes, and Upgrade Guide (28 minute read)
Apache Iceberg 1.12.0 broadens production support for major v3 capabilities, including Variant, geospatial types, row lineage, deletion vectors, encrypted metadata, and REST catalog operations. However, it also removes platform support and deprecated APIs, so treat this as a substantial upgrade and test compatibility carefully before rollout.
The Inference Gap (22 minute read)
A six-week analysis of 43,261 Fable 5 calls found that the model name stayed the same while the delivered reasoning budget varied sharply, with many invocations showing no thinking at all. The broader warning is an inference gap: benchmark performance may depend on far more sequential reasoning than production users actually receive.
Microsoft, Google back Apache Ossie to make enterprise data and AI platforms more interoperable (4 minute read)
Microsoft and Google are backing Apache Ossie, an open semantic-model interchange specification supported by 60+ data and AI vendors. Using standard JSON/YAML schemas, Ossie enables portable dataset, metric, and relationship definitions across disparate BI platforms, eliminating redundant semantic modeling and metric drift.
Why Your AI Agent Pipeline Is Slow (And How to Fix It Without Changing Models) (5 minute read)
A production AWS Bedrock guardrail pipeline for AI agents reduced action latency from 13,874 ms to 1,824 ms without altering models or accuracy. The bottleneck was execution topology: independent safety checks ran sequentially. By parallelizing independent layers and short-circuiting cheap reject rules early, the team dramatically reduced execution latency and compute costs.
Faster String Aggregations with Dimension Tables (16 minute read)
String-heavy aggregations waste time hashing and comparing the same text values repeatedly. DuckDB shows that replacing those columns with small integer surrogate keys through dimension tables can make grouping cheaper, reduce memory pressure, and lower spill risk. The broader takeaway: normalize low-cardinality strings before large analytical aggregations when performance matters.
Curated deep dives, tools and trends in big data, data science and data engineering 📊
Join 590,000 readers for
one daily email