» Tag
big-data
5 postsHow Grab Scaled Its Data Lake by Adopting Apache Iceberg
Grab's migration from Hive Parquet to Apache Iceberg: performance gains, cost savings, and the open-sourced UnifiedSparkCatalog for Spark.
Apache Parquet in 2026: The Quiet Format Enters Its Loudest Decade
Thirteen-year-old Apache Parquet faces its busiest development era yet, driven by lakehouse and AI demands: variant types, geospatial support, and footer debates.
Apache Iceberg's Variant Type: How Shredding Speeds Up JSON Analytics
Apache Iceberg v3's Variant type and shredding technique make JSON data both flexible and fast to query. Here's how it works under the hood.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comIndexing the Data Lake for Online Point Queries
Spotify and others leverage RAP to optimize point queries in data lakes.
Advanced Partitioning Strategies for Petabyte-Scale Tables
Learn about effective partitioning strategies for petabyte-scale datasets. Best practices for performance and management.