Skip to content
← Technical Diagram Library
Data Engineering

Spark Lakehouse Data Flow

Connect batch and streaming ingestion to Spark processing, lakehouse storage and analytical consumption.

Apache SparkPySparkKafkaIcebergDelta LakeTrinoAirflow

Spark Lakehouse Data Flow

Modern data platforms converge batch and streaming into governed storage. Spark sits inside a wider control plane: orchestration, table formats, query engines and observability all influence production behavior.

01

Spark is part of a platform, not the entire platform.

02

Table format and query-engine choices affect operational design.

03

Orchestration and observability belong in the architecture.

Spark Lakehouse Data FlowConnect batch and streaming ingestion to Spark processing, lakehouse storage and analytical consumption.01SourcesDB • files • events02Kafka / Object StoreIngestion03Spark / PySparkBatch + streaming04Iceberg / DeltaLakehouse tables05Trino / BIConsumption06Airflow + ObservabilitySchedule • SLA • metrics
How to read it

Follow the handoffs, then ask where evidence exists.

01

Sources

DB • files • events

02

Kafka / Object Store

Ingestion

03

Spark / PySpark

Batch + streaming

04

Iceberg / Delta

Lakehouse tables

05

Trino / BI

Consumption

06

Airflow + Observability

Schedule • SLA • metrics

Architecture questions
Spark is part of a platform, not the entire platform.
Table format and query-engine choices affect operational design.
Orchestration and observability belong in the architecture.
Technology context
Apache SparkPySparkKafkaIcebergDelta LakeTrinoAirflow

The diagram is intentionally architectural rather than vendor-specific. Use it as a mental model, then map the components to the actual environment.

Need the architecture applied?

Use the visual model as the starting point for a workshop or technical review.