@apachehudi

Official twitter handle of Apache Hudi, an open data lakehouse platform. https://nitter.cf/t.co/SXay7oHNah

Joined January 2019
Hudi 1.0 is the most powerful release to date for data lakehouses. Read the blog for details: Secondary Indexing, Expression Indexes, Partial Updates, Non-blocking Concurrency Control, New LSM timeline, +more: hudi.apache.org/blog/2024/12… #datalakehouse #opentableformat
10
36
2,898
Apache XTable (Incubating) is how a Hudi table becomes an Iceberg table without rewrites. Same data files on object storage. Two metadata layers describing them — one Hudi, one Iceberg — each consumable by a different engine ecosystem.
1
2
141
The interop pattern that lets the engineering decision be "write to the format with the best primitives for our workload" — not "write to what our downstream engines understand." Links ↓ #ApacheHudi #ApacheXTable
1
1
22
The layout of a Hudi table is what makes everything else work. Hierarchy, top to bottom ↓
1
1
4
211
Storage format is versioned. Hudi 1.x ships backward-compatible writes — upgrade readers first, then writers in any order. No coordinated downtime. Links ↓ #ApacheHudi #StorageDesign
1
23
Tencent Cloud EMR added native Apache Hudi support in v2.2.0 as a first-class BigData component. Native Hudi on managed Spark + Flink. Data on HDFS, COS, or CHDFS ↓
1
1
170
Open lakehouse is cloud-portable in a way closed proprietary alternatives can't match — tables move across clouds without rewrites. Links ↓ #ApacheHudi #CloudPortability
1
27
Tencent Cloud EMR: intl.cloud.tencent.com/produ… Apache Hudi powered-by list (cloud integrations): hudi.apache.org/powered-by/
28
Listing partitions on S3 doesn't scale. Hudi's metadata table fixes this. Cloud object storage LIST calls are slow — seconds on small tables, many minutes on partitioned tables with thousands of folders. Every reader and writer pays.
1
1
3
155
Big Metadata paper (Google, related concept): vldb.org/pvldb/vol14/p3083-e…
23