Data engineering ยท science ยท architecture

IngestThis

Working notes for people who move data for a living. Pipelines, lakehouses, table formats, and the architecture underneath them.

Articles
550+
Topics
Iceberg, pipelines, AI
Submissions
Open
Price
Free
  1. 01Why the Iceberg DataFusion Integration Is Moving to Apache DataFusionWhy the Iceberg DataFusion integration moved to the DataFusion project, and what the split means for users, Comet, and iceberg-rust contributors.2026-09-21
  2. 02What Iceberg v4's Proposed FILE Type Means for Multimodal TablesIceberg v4's proposed FILE type brings first-class media references to tables, via Parquet's FILE logical type, ranges, checksums, and pre-signed URLs.2026-09-21
  3. 03Fast Classification Models, LLMs, and the Apache Iceberg LakehouseHow fast classification models like Jev alongside open alternatives such as GLiClass compare with LLMs, and how to run both together inside an Apache Iceberg lakehouse.2026-09-21
  4. 04How Apache Ossie Is Deciding What Agents and BI Tools Can Ask a Semantic LayerHow Apache Ossie's layered query design gives AI agents both a constrained dimensional interface and a grain-safe SQL interface for semantic layers.2026-09-21
  5. 05CVE-2026-73334 and the Trust Boundary Inside an Encrypted Parquet FileCVE-2026-73334 lets a tampered Parquet footer route a reader's KMS token to an attacker. Here's the fix, Iceberg's safe path, and how to audit your lakehouse.2026-09-21
  6. 06Parquet Page Indexes and the Last Mile of Pruning in Apache IcebergParquet page indexes can cut selective Iceberg scans by an order of magnitude on sorted data. How they work, what they cost, and how to lay out tables.2026-09-21
  7. 07How Apache Polaris Plans to Share Iceberg Tables Across OrganizationsApache Polaris's Open Sharing proposal adds first-class shares, external consumers, and listings so any Iceberg REST engine can read shared tables.2026-09-21
  8. 08Keeping Lakehouse Traffic Off the Public Internet With Apache PolarisA private Polaris lakehouse still leaks data if the storage hop goes public. How to bind vended credentials to private networks on AWS, Azure, and GCP.2026-09-21

Reference shelf

Long reads from around the web
Semantic Layer

The Semantic Layer: Definitive Guide

A comprehensive guide to the Semantic Layer โ€” how it creates a single source of truth for metrics, powers headless BI, and makes AI agents answer business questions accurately.

Read
Apache Polaris

Apache Polaris: The Catalog Standard for Lakehouses and AI

How Apache Polaris is emerging as the universal Iceberg catalog standard, enabling multi-engine interoperability and governed AI access across the lakehouse ecosystem.

Read
Table Formats

What Are Table Formats and Why Were They Needed?

The origin story of open table formats โ€” the problems with Hive, why Apache Iceberg, Delta Lake, and Hudi were created, and what they unlock for modern data platforms.

Read
Dremio

What Is Dremio?

A clear-eyed breakdown of what Dremio is, how its semantic layer, query federation, Reflections, and Apache Arrow Flight power the Intelligent Lakehouse Platform.

Read
Apache Iceberg

What Apache Iceberg Native Actually Means

Not all 'Iceberg support' is equal. This piece breaks down what it means to be genuinely Apache Iceberg native versus bolt-on, and why it matters for your lakehouse.

Read
Open Source

Open Source and the Data Lakehouse

How the Apache Software Foundation's open-source projects โ€” Iceberg, Arrow, Parquet, Polaris โ€” form the modular foundation of the modern open data lakehouse.

Read
Agentic AI

What Is Agentic Analytics?

Agentic AI is reshaping how organizations interact with data. This guide explains agentic analytics, the role of the semantic layer, and why query performance matters for AI agents.

Read
Data Lakehouse

Definitive Guide to the Data Lakehouse

The complete, authoritative guide to the Data Lakehouse architecture โ€” what it is, why it supersedes the data warehouse + data lake combination, and how to build one.

Read
AI & Performance

How Dremio Keeps Agentic Analytics Fast Without Manual Tuning

How Dremio's layered autonomous performance architecture โ€” Reflections, caching, vectorized execution โ€” handles unpredictable AI agent query patterns at interactive speed.

Read

Elsewhere

Pitch an idea: alex@ingestthis.com