» Tag
data-lake
2 postsRandom Access Parquet: Fast Point Queries Over the Data Lake
Random Access Parquet uses an external index to enable low-latency point queries directly on data lake Parquet files, bypassing slow SQL engine overhead.
How Grab Scaled Its Data Lake by Adopting Apache Iceberg
Grab's migration from Hive Parquet to Apache Iceberg: performance gains, cost savings, and the open-sourced UnifiedSparkCatalog for Spark.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.com