Tag: Apache Iceberg
All the articles with the tag "Apache Iceberg".
- 31 MIN READ•Aug 25, 2026
Agent-Driven Storage Tiering for Apache Iceberg: Moving Cold Data Without Breaking Queries
A background agent can move cold Iceberg partitions to cheaper tiers without breaking live queries. Heatmaps, path-safe moves, and restore paths.
Apache Icebergstorage tieringcost optimization - 31 MIN READ•Aug 25, 2026
DataFusion Comet 1.0 and What Native Rust Scans Change for Spark on Iceberg
DataFusion Comet 1.0 replaces Spark Iceberg scans with native Rust. What speeds up, what still falls back to the JVM, and how to deploy it.
Apache IcebergApache SparkDataFusion Comet - 31 MIN READ•Aug 25, 2026
High-Throughput Branch Merging: Automating Concurrency and Conflict Resolution in Multi-Branch Iceberg Pipelines
High-throughput Iceberg branch merges need conflict detection and automation. How to reconcile concurrent writes without stalling pipelines.
Apache Icebergbranchesconcurrency - 31 MIN READ•Aug 25, 2026
Multi-Cloud REST Catalog Topologies: Running Apache Polaris Across AWS, Azure, and GCP
Polaris can catalog Iceberg tables across AWS, Azure, and GCP. Four topologies, credential vending, and the tradeoffs of each design.
Apache PolarisREST catalogmulti-cloud - 32 MIN READ•Aug 25, 2026
Parquet-Only Manifests in Iceberg v4: Why the Metadata Layer Is Going Columnar
Iceberg v4 is moving manifests from Avro to Parquet so planners can read only the stats they need. Why the metadata layer is going columnar.
Apache IcebergIceberg v4Parquet - 31 MIN READ•Aug 25, 2026
Semantic Layer Federation: One Logical Model Over Data on Three Clouds
One logical model over Iceberg and databases on three clouds. Pushdown, egress, Reflections, and where semantic federation still breaks.
semantic layerfederationmulti-cloud - 31 MIN READ•Aug 25, 2026
Serverless Iceberg Ingestion with PyIceberg and DuckDB: Micro-Batches Without a Spark Cluster
Land small Iceberg micro-batches with PyIceberg and DuckDB in a serverless function. Commits, concurrency, and why Spark is the wrong default.
Apache IcebergPyIcebergDuckDB - 31 MIN READ•Aug 25, 2026
Zero-Copy Warehouse Modernization: Migrating Legacy Databases to Apache Iceberg Without Downtime
Move a legacy warehouse to Iceberg without downtime by virtualizing first. Consumer cutover, parity checks, and background copy without double-ETL.
Apache Icebergmigrationfederation - 31 MIN READ•Aug 24, 2026
The Hidden Cost of Tiny Iceberg Commits
Trace what one tiny Iceberg commit writes, then model hourly, per-minute, and per-second cadences so streaming costs become arithmetic, not adjectives.
Apache Icebergstreamingmetadata - 31 MIN READ•Aug 24, 2026
Deletion Vectors vs Position Deletes vs Equality Deletes: The Iceberg Delete Story in 2026
Position deletes, equality deletes, and deletion vectors compared from the Iceberg spec: what each writes, how readers apply it, and when to use which.
Apache Icebergdeletion vectorsposition deletes - 31 MIN READ•Aug 24, 2026
Iceberg Is Becoming a Library, Not Just a Table Format
Iceberg is turning from a JVM table format into a library other systems embed. What that shift changes for engines, catalogs, and the spec itself.
Apache Iceberglibrariesecosystem - 31 MIN READ•Aug 24, 2026
Iceberg Is Escaping the JVM: Why Rust, Go, Python and C++ Implementations Matter
Rust, Go, Python, and C++ Iceberg implementations change who can write the format. Why multi-language clients matter more than another JVM engine.
Apache IcebergRustPython