Enroll now in our Apache Iceberg Course to gain hands-on experience through live interactive sessions, real-world projects, and personalized mentorship. The course is fully aligned with modern data lakehouse and open data ecosystem best practices, enabling you to confidently build, manage, and optimize Iceberg-based data platforms.
Prerequisites
- Basic understanding of databases and SQL
- Knowledge of data engineering concepts
- Familiarity with cloud storage or the Hadoop ecosystem
- Basic knowledge of programming or scripting
What Will You Learn
- What is a “data lakehouse” and why it matters
- What is a “table format” in the lake context
- Overview of Iceberg’s architecture: table metadata, manifest lists, data files, catalog service, layers
- Support for ACID transactions, schema & partition evolution
- Installation, dependencies & environment (on-prem, cloud object storage: S3/MinIO, etc.)
- Integrating with compute engines (Apache Spark, Apache Flink, SQL engines such as Hive/Trino/Presto)
- Creating tables: syntax, options, choosing formats (Parquet, ORC, Avro)
- Table maintenance: compaction, expiring older snapshots, garbage collection/clean-up
- Schema evolution (add/drop/rename fields) and partition evolution
- Upserts, deletes, merges (row-level operations)
- Optimizing table layout (file sizing, partitioning strategies)
- Auditing and metrics: measuring performance, monitoring usage, and setting alerts
- CDC support: how Iceberg handles change streams, late data, and merges