Flowchart examples · Architecture

ETL data pipeline flowchart

Scheduled extract to refreshed dashboards, with validation gates and alerting when data is bad.

ETL data pipeline flowchart

ETL data pipeline flowchartNoYesYesNoScheduled run startsExtract from databasesand APIsLand raw files instorageSchema and rowcounts valid?Alert data team, stoprunTransform: clean, join,aggregateLoad into warehousetablesData quality testspass?Quarantine batch, keeplast good dataRefresh dashboardsRun complete

A worked example to adapt. Rename steps, add branches and owners to match how your team actually works.

Edit the steps to match your process. The free preview (no sign-up) drafts up to 8 main steps; add branches and the remaining steps in the editor.
Mermaid source for this diagram
flowchart TD
  s(["Scheduled run starts"])
  ext["Extract from databases and APIs"]
  raw[("Land raw files in storage")]
  d1{"Schema and row counts valid?"}
  alert["Alert data team, stop run"]
  tr["Transform: clean, join, aggregate"]
  load["Load into warehouse tables"]
  d2{"Data quality tests pass?"}
  quar["Quarantine batch, keep last good data"]
  bi["Refresh dashboards"]
  e(["Run complete"])
  s --> ext
  ext --> raw
  raw --> d1
  d1 -->|No| alert
  d1 -->|Yes| tr
  tr --> load
  load --> d2
  d2 -->|Yes| bi
  d2 -->|No| quar
  quar --> alert
  bi --> e

Paste into any Markdown tool that renders Mermaid, such as GitHub.

About this etl data pipeline flowchart

A data pipeline diagram is most useful when it shows where bad data is caught. This example adds two gates: a schema and volume check on raw data, and quality tests after loading, so broken data never reaches dashboards.

Many modern stacks load first and transform inside the warehouse (ELT); the same gates apply.

Step by step

  1. Extract. Pull from source databases and APIs on a schedule.
  2. Land and validate. Store raw data, then check schema and row counts.
  3. Transform. Clean, join and aggregate into analysis-ready tables.
  4. Load and test. Load to the warehouse and run data quality tests.
  5. Serve. Refresh dashboards only with data that passed.

How to make it in flow-chart.io

  1. Start from the example. Edit the text in the generator box above so it names your own sources, orchestrator and warehouse, then press Generate. The free preview needs no sign-up.
  2. Add the branches. Add the decision points that matter for you, for example late-arriving data? or backfill run. Each decision becomes a diamond with labeled outcomes.
  3. Assign owners. Save the diagram to the editor (free account) and rename steps to show who does what: orchestrator, warehouse and BI tool.
  4. Share or export. Share a read-only view link: people with the link can view the diagram but not edit it. Downloads as PNG, SVG or PDF are part of Flow Pro (see pricing).

Tips

Frequently asked questions

What is the difference between ETL and ELT?
ETL transforms data before loading it into the warehouse; ELT loads raw data first and transforms it inside the warehouse.
Where should data quality checks go?
At least after extraction (schema, volume) and after transformation (business rules), as in this example.
Can I generate a diagram for my stack?
Yes. Name your tools in the generator box, for example Airflow, dbt and Snowflake, and press Generate.

Related

AWS architecture for a web appCI/CD pipeline flowchartIncident response process flowchartData pipeline diagramsData flow diagramsAll flowchart examples

Build it from your own words

Describe your process and get an editable diagram. Free to build and edit; Flow Pro adds downloads and unlimited projects.

Start free See pricing