Big Data & Distributed Systems
Build the mental model behind modern data platforms.
- Scale-up vs scale-out
- Partitioning and parallelism
- Distributed storage
- Failure and retries
- Batch vs streaming
Understand distributed data systems and build practical batch, streaming and lakehouse pipelines with production-oriented performance and operability.
The exact sequence is adjusted to the audience. Foundation topics can be compressed for experienced teams; architecture and troubleshooting can be expanded for advanced programs.
Build the mental model behind modern data platforms.
Learn Spark from execution model to production tuning.
Design event-driven data movement.
Query data across platforms efficiently.
Connect object storage, table formats and compute engines.
Make pipelines observable and repeatable.
Exercises emphasize observation, implementation, failure and diagnosis rather than command copying.
Build and tune a PySpark batch pipeline
Inspect a Spark execution plan and shuffle
Process Kafka events using Structured Streaming
Query lakehouse data with Trino
Orchestrate a multi-stage pipeline
Programs can focus on Spark performance, Kafka streaming, Trino, lakehouse architecture or an end-to-end data-platform journey.