D · The Data Platform on One Page

This page puts Steep’s whole data platform on one map, and then names each piece in open-source, AWS and Alibaba Cloud terms. Use the map to see where a tool sits; use the table to turn a job ad’s list of products back into the chapters of this book.

An illustrated map of a small tea town seen from slightly above, on cream paper. Cobblestone paths join seven buildings in a row from left to right. First, an open kitchen with a tomato-red roof: blank paper tickets hang from a long metal rail under its roof, above stoves with steaming pots. Second, a small pantry with a navy roof, its open front showing tall shelves of jars. Third, four identical small kitchens side by side under four red roofs, each with a pot on a counter; the window of the third one glows mustard yellow. Fourth, a tea factory with a tall brick chimney, where a conveyor belt carries tea leaves past a metal tank and stacked crates. Fifth, a round tea bar under a teal canopy, with teacups riding a belt in a loop. Last, on the shore of a calm teal lake, a small boathouse with a teal roof and a rowing boat at its dock, and behind it a tall warehouse with open doors and rows of shelves full of boxes. A thin tomato-red path enters at the lower left, runs through every building, and leaves at the upper right.

The town of Part II. From the left: the kitchen with the ticket rail (Kafka, Chapter 8); the pantry with its shelves of jars (storage, Chapter 9); four small kitchens side by side, one lit brighter than the rest (Hive and Spark, Chapter 10); the tea factory (batch pipelines, Chapter 13); the conveyor-belt tea bar (streaming, Chapter 14); the boathouse on the lake (the lakehouse, Chapter 11); and the warehouse with its shelves (Chapter 12). The red path is the road of one order.

The map

Each box names one job and the chapter that teaches it. On a computer, the map reads from left to right; on a phone, from top to bottom. The list under the map links to every chapter.

Data lakeWarehouse layers (Chapter 12)Warehouse layersCh 12Apps and web (Chapter 7)Apps and webevents from iOS,Android and webCh 7Orders database (Chapter 7)Orders databaseMySQL: one rowper orderCh 7, 9Kafka (Chapter 8)Kafkaevents and changes,kept in partitionsCh 8Change data capture (CDC) (Chapter 7)Change datacapture (CDC)reads the binlogCh 7Stream processing (Chapter 14)StreamprocessingFlink: answersin secondsCh 14Batch compute (Chapter 10)Batch computeSpark (or Hive):nightly jobsCh 10Scheduler (Chapter 13)SchedulerAirflow,DolphinSchedulerCh 13Live alerts (Chapter 14)Live alertsfresh numbers fromthe streamCh 14Dashboards (Chapter 5)Dashboardscharts for peoplewho decideCh 5Query engines (Chapter 11)Query enginesTrino, ClickHouse,Doris, StarRocksCh 11Data quality tests, reconciliation, contracts and lineage watch every step (Chapter 15)✓ Data quality tests, reconciliation, contracts and lineage watch every stepCh 15Catalog (Chapter 10)CatalogHive metastoreCh 10Table format (Chapter 11)Table formatnot yet at Steep:Iceberg, Delta Lake…Ch 11Files (Chapter 9)FilesParquet, indt= partitionsCh 9Storage (Chapter 10)StorageHDFS orobject storageCh 10ODSraw copiesDWDclean detailsDWSsummariesADSone per reportSOURCESCAPTURE & CARRYSTORECOMPUTEMODELSERVEnightly copy (Ch 13)The road of order A1024, from the app to the warehouse summaries (DWS)Where A1024 is missing: the app never sent its order_completed event, and the CEO dashboard's ADS table counts those eventsStarts jobs (no data flows here)Data lakeWarehouse layers (Chapter 12)Warehouse layersCh 12Apps and web (Chapter 7)Apps and webevents from iOS,Android and webCh 7Orders database (Chapter 7)Orders databaseMySQL: one rowper orderCh 7, 9Kafka (Chapter 8)Kafkaevents and changes,kept in partitionsCh 8Change data capture (CDC) (Chapter 7)Change datacapture (CDC)reads the binlogCh 7Stream processing (Chapter 14)StreamprocessingFlink: answersin secondsCh 14Live alerts (Chapter 14)Live alertsfresh numbers fromthe streamCh 14Batch compute (Chapter 10)Batch computeSpark (or Hive):nightly jobsCh 10Scheduler (Chapter 13)SchedulerAirflow,DolphinSchedulerCh 13Query engines (Chapter 11)Query enginesTrino, ClickHouse,Doris, StarRocksCh 11Dashboards (Chapter 5)Dashboardscharts for peoplewho decideCh 5Data quality tests, reconciliation, (Chapter 15)✓ Data quality tests, reconciliation,contracts and lineage watch everystepCh 15Catalog (Chapter 10)CatalogHive metastoreCh 10Table format (Chapter 11)Table formatnot yet at Steep:Iceberg, Delta Lake…Ch 11Files (Chapter 9)FilesParquet, indt= partitionsCh 9Storage (Chapter 10)StorageHDFS orobject storageCh 10ODSraw copiesDWDclean detailsDWSsummariesADSone per reportSOURCESCAPTURE & CARRYSTORECOMPUTESERVECOMPUTEMODELSERVEnightly copy (Ch 13)The road of order A1024, from the app tothe warehouse summaries (DWS)Where A1024 is missing: the app neversent its order_completed event, and theCEO dashboard's ADS table counts thoseeventsStarts jobs (no data flows here)

Steep’s data platform. The tomato line follows order A1024, Mia’s oolong milk tea, from her phone to the warehouse summaries (DWS). The CEO dashboard counts events, so the order never reached it.

Reading the map: the road of one order

  1. Apps and web (Chapter 7). Mia paid in the iOS app. The app sent its events, but not order_completed: that one event was lost here, in the app itself (Chapter 15).
  2. The orders database (Chapter 7, Chapter 9). The order service saved order A1024 as one row in MySQL. For orders and money, this is the system of record.
  3. The nightly copy and change data capture (Chapter 13, Chapter 7). Once a night, Steep copies the orders table into the lake: that is how the row of A1024 reached the lake. MySQL also wrote every change to its binlog: the insert at payment, then the update to completed. A CDC tool such as Debezium reads the binlog and passes each change to Kafka, and Steep keeps a copy of the binlog in the lake (cdc_orders_binlog).
  4. Kafka (Chapter 8). Events and changes wait in order, in partitions, for a set time (seven days by default), whether or not anyone has read them. Each reader keeps its own place, its offset.
  5. The data lake (Chapter 9, Chapter 10, Chapter 11). Parquet files in dt= folders, on HDFS or object storage. A catalog gives the folders table names. Steep’s tables are plain Hive tables; a table format, drawn dashed, would add commits, updates and time travel.
  6. Batch compute and the scheduler (Chapter 10, Chapter 13). Every night, the scheduler starts Spark jobs in the order of their dependencies, one partition per day.
  7. The warehouse layers (Chapter 12). ODS keeps raw copies, DWD cleans them, DWS summarises, and ADS serves one report each. A1024 is in DWD and DWS. The CEO dashboard’s ADS table counts order_completed events, so the order stops there.
  8. Query engines and dashboards (Chapter 11, Chapter 5). The CEO dashboard never counted A1024. Reports built on the orders, such as finance’s, did.
  9. The stream path (Chapter 14). Flink reads the same Kafka topics and answers within seconds: alerts now, exact numbers later from the batch path.
  10. Quality and lineage (Chapter 15). Tests compare the steps with each other, every day, and lineage says what else a broken table touches.

The same jobs, by their product names

Job ads list products, not jobs. Find the product in the table, read the job it does, and go to the chapter. Names change, and a product can be retired; the job matters more than the brand. For example, Amazon QuickSight was renamed Amazon Quick Sight (two words), a feature of Amazon Quick, and MaxCompute was once called ODPS. Products marked “closest match” do a similar job in a different way.

Job What it does Open source AWS Alibaba Cloud Ch
Event collection Records what people do in an app or on a website, and sends it to a server Snowplow; RudderStack ‡ No single product: Guidance for Clickstream Analytics on AWS (closest match) Quick Tracking; Simple Log Service (SLS) web tracking 7
Change data capture Reads a database’s change log (the binlog) and passes on every change Debezium; Apache Flink CDC; Canal AWS Database Migration Service (AWS DMS) Data Transmission Service (DTS) 7
Data integration and loading Copies data from one system into another on a schedule or all the time, such as the nightly copy of the orders table, or Kafka into the lake Kafka Connect; DataX; Apache SeaTunnel AWS Glue; Amazon Data Firehose DataWorks Data Integration 8, 13
Message queue Keeps events in order, so that many readers can read them, each at its own pace Apache Kafka Amazon Managed Streaming for Apache Kafka (Amazon MSK); Amazon Kinesis Data Streams (closest match) ApsaraMQ for Kafka; DataHub (closest match) 8
Stream processing Computes results from events as they arrive, within seconds Apache Flink Amazon Managed Service for Apache Flink Realtime Compute for Apache Flink 14
Storage Keeps the files, cheaply and with several copies HDFS Amazon S3 Object Storage Service (OSS); OSS-HDFS 9, 10
Table format Turns folders of files into tables with commits, updates and time travel Apache Iceberg; Delta Lake; Apache Hudi; Apache Paimon Amazon S3 Tables (Iceberg) Data Lake Formation (DLF), with Paimon tables 11
Metadata catalog Knows each table’s columns, files and partitions, so that engines can find them Hive Metastore; Apache Polaris; Unity Catalog; Apache Gravitino AWS Glue Data Catalog Data Lake Formation (DLF) 10, 11
Batch compute Runs big jobs over many machines, usually once a night Apache Spark; Apache Hive Amazon EMR; AWS Glue E-MapReduce (EMR); MaxCompute 10
SQL transformations Builds the warehouse layers from SQL files, in the order of their dependencies, with tests dbt (open-source dbt Core) No product of its own: dbt on Amazon Redshift or Amazon Athena DataWorks with MaxCompute SQL 12, 15
SQL on files Runs SQL on files where they lie, and keeps no data of its own Trino; Presto Amazon Athena E-MapReduce (EMR) with Trino 11
Scheduler Starts jobs in the order of their dependencies, on time, and retries failures Apache Airflow; Apache DolphinScheduler Amazon Managed Workflows for Apache Airflow (Amazon MWAA) DataWorks 13
Warehouse and OLAP database Keeps its own copy of the data and answers analysis queries fast ClickHouse; Apache Doris; StarRocks Amazon Redshift MaxCompute; Hologres; AnalyticDB for MySQL; EMR Serverless StarRocks 11, 12
BI dashboards Charts and dashboards for the people who decide Apache Superset; Metabase Amazon Quick Sight Quick BI 5
Data quality Tests the data every day and alerts an owner when a test fails dbt data tests; Great Expectations (GX Core); Deequ AWS Glue Data Quality DataWorks Data Quality 15
Lineage and metadata Shows which table is built from which, what it means and who owns it OpenLineage (with Marquez); DataHub; Apache Atlas Amazon SageMaker Catalog (built on Amazon DataZone) DataWorks Data Map 12, 15

Notes on the table:

  • Closest match. Amazon Kinesis Data Streams and Alibaba Cloud DataHub are streaming services of their own, not Kafka. AWS has no single event-collection product: its Guidance for Clickstream Analytics on AWS is a ready-made design with app SDKs: small code libraries inside the app that send its events into AWS services.
  • ‡ Source-available, not open source. Snowplow’s trackers are open source, but its collector and pipeline now use the Snowplow Limited Use License. RudderStack’s server uses the Elastic License 2.0, and so does Soda Core, a data-quality tool. Read the licence before you build on any of them.
  • One product, several jobs. DataWorks loads data, runs SQL jobs on a schedule, tests data quality and keeps lineage. MaxCompute is both the warehouse and its compute engine. Data Lake Formation is a catalog and managed table storage. DataX is the open-source version of DataWorks’ data integration. In real platforms, one product often does several of these jobs. The map shows the jobs, not the products that vendors sell.
  • Steep’s own stack. Steep uses the open-source pieces named in the chapters: MySQL, Kafka, Parquet files with a Hive metastore (plain Hive tables, no table format yet; Chapter 11), Spark and a scheduler. DuckDB is the engine inside this book.