Skip to content
All training programs
INGEST → PROCESS → STREAM → ORCHESTRATE → SERVE

Data Engineering & Big Data

Understand distributed data systems and build practical batch, streaming and lakehouse pipelines with production-oriented performance and operability.

Typical duration3–8 days; modular Spark, Kafka or platform-specific tracks available
DeliveryOnline · On-site · Hybrid · Data-platform lab
Practical work5 representative labs
Outcomes
Explain distributed storage and compute trade-offs
Build and tune Spark / PySpark workloads
Design batch and streaming pipelines
Understand lakehouse, orchestration and distributed SQL patterns
Audience & prerequisites
Who it is forData engineersPlatform engineers supporting data workloadsDevelopers moving into distributed dataArchitects evaluating modern data platforms
PrerequisitesLinux fundamentalsBasic Python and SQLDistributed-systems concepts are introduced as needed
Representative curriculum

Training topics, organized as engineering modules.

The exact sequence is adjusted to the audience. Foundation topics can be compressed for experienced teams; architecture and troubleshooting can be expanded for advanced programs.

01

Big Data & Distributed Systems

Build the mental model behind modern data platforms.

  • Scale-up vs scale-out
  • Partitioning and parallelism
  • Distributed storage
  • Failure and retries
  • Batch vs streaming
02

Apache Spark

Learn Spark from execution model to production tuning.

  • Driver, executors and cluster managers
  • DataFrames and Spark SQL
  • PySpark
  • Partitions and shuffles
  • Caching and joins
  • Performance diagnostics
03

Streaming & Kafka

Design event-driven data movement.

  • Kafka brokers, topics and partitions
  • Producers and consumers
  • Offsets and delivery semantics
  • Spark Structured Streaming
  • Reliability and back pressure
04

Distributed SQL & Serving

Query data across platforms efficiently.

  • Trino architecture
  • Connectors and catalogs
  • Query planning
  • Federated query
  • Performance and security
05

Lakehouse Architecture

Connect object storage, table formats and compute engines.

  • Data lakes vs warehouses vs lakehouse
  • Iceberg / Delta concepts
  • Schema evolution
  • Partition design
  • Spark / Trino interoperability
06

Orchestration & Operations

Make pipelines observable and repeatable.

  • Airflow concepts
  • Dependency management
  • Retries and idempotency
  • Data-quality checks
  • Monitoring and troubleshooting
Hands-on work

Labs are part of the learning path.

Exercises emphasize observation, implementation, failure and diagnosis rather than command copying.

01

Build and tune a PySpark batch pipeline

02

Inspect a Spark execution plan and shuffle

03

Process Kafka events using Structured Streaming

04

Query lakehouse data with Trino

05

Orchestrate a multi-stage pipeline

Customization

Align the program to your engineering environment.

Programs can focus on Spark performance, Kafka streaming, Trino, lakehouse architecture or an end-to-end data-platform journey.

Apache SparkPySparkKafkaHadoop / HDFSFlinkTrinoAirflowIcebergDelta LakeLakehouse
Discuss this program