How do declarative pipeline definitions and workflow orchestration fit together?
Summary
- They are two complementary layers, not competing choices. A declarative pipeline defines the relationships between datasets (the what); workflow orchestration defines the relationships between tasks (the when and how).
- In Databricks Lakeflow, Lakeflow Pipelines is the declarative layer and Lakeflow Jobs is the orchestration layer. A pipeline runs as a single task inside a job. Both are generally available.
- A declarative pipeline resolves its own dependency graph, runs datasets in the right order and in parallel, processes only new or changed data, and retries failures, so you describe the result rather than hand-wire the steps.
- Orchestration adds what a single pipeline does not: schedules and event triggers, conditional branches and loops, per-task retries and alerts, and one view across notebooks, SQL, dbt, and pipelines.
- The clean division: put dataset logic in the pipeline and let the engine manage it; put branching, loops, and heterogeneous steps in the job around it.
How do declarative pipeline definitions and workflow orchestration fit together?
They are two complementary layers, not competing choices. A declarative pipeline defines the relationships between datasets: you describe the tables you want, and the engine works out the order to build them, processes only what changed, and retries failures on its own. Workflow orchestration defines the relationships between tasks: it decides when work runs, in what sequence, under which conditions, and what happens on success or failure. In Databricks Lakeflow, Lakeflow Pipelines is the declarative layer and Lakeflow Jobs is the orchestration layer, and a pipeline runs as a single task inside a job. Both are generally available.
What is the difference between a declarative pipeline and orchestration?
A declarative pipeline is data-centric. You declare target datasets, streaming tables for incremental ingestion and materialized views for derived tables, and attach data quality expectations. The engine then resolves the dependency graph between those datasets, runs them in the right order and in parallel where it can, processes new or changed data incrementally, and retries transient failures progressively, from the Spark task up to the whole pipeline. You do not hand-wire the order or write per-record loops. You describe the result.
Orchestration is task-centric. A job is one or more tasks arranged as a directed acyclic graph, where each task is a unit of work: a notebook, a SQL query, a Python script, a dashboard refresh, a dbt project, or a pipeline. Jobs add what a single pipeline does not: schedules and event triggers, conditional branches, loops over a set of inputs, per-task retries and notifications, and unified monitoring across every task type.
Which layer handles which concern
The two divide the work cleanly:
| Concern | Declarative pipeline | Workflow orchestration |
|---|---|---|
| Dependencies between datasets | Resolved automatically from the definitions | Not its job |
| Order and parallelism inside the data flow | Handled by the engine | Not its job |
| Incremental processing and data quality | Built in, through streaming tables, materialized views, and expectations | Not its job |
| When and how often work runs | Set by the job that runs the pipeline | Schedules, event triggers, continuous mode |
| Conditional branches and loops | Not expressed here | If/else and For each tasks |
| Sequencing heterogeneous steps (notebooks, SQL, dbt) | Not its job | Coordinated as tasks in the job |
| Retries and alerting | Automatic within the pipeline | Per task, across the whole job |
How Databricks approaches this
In Lakeflow, the two layers are built to compose. A pipeline integrates into a job as a pipeline task. How it runs follows the job: in a triggered or scheduled job, the pipeline task starts a single update and stops when it finishes; in a continuous job, the pipeline task keeps the pipeline running. You can add a pipeline as a task from the Jobs UI, the Lakeflow Pipelines UI, or in SQL.
That composition is where the model pays off. The pipeline owns everything about the data: it analyzes dependencies between datasets, orchestrates their order and parallelism, keeps materialized views current by reprocessing only new or changed source data, and enforces data quality inline. See the declarative pipelines concepts. The job owns everything around the data: it can run a preprocessing notebook, branch on a condition with an If/else task, run the pipeline, then refresh a dashboard, all with per-task retries and one view of the run. Because both are part of Lakeflow and governed by Unity Catalog, you get one lineage graph and one place to monitor, whether a step is a declarative pipeline or a scripted task. See Announcing the general availability of Databricks Lakeflow.
Patterns
- Pipeline as a task: embed a declarative pipeline in a broader job when you need steps around it, such as a conditional preprocess, then the pipeline, then post-processing and a dashboard refresh.
- Orchestrator first: put the branching, loops, and heterogeneous steps in the job, and call a pipeline only where declarative table logic makes sense.
- Continuous data flow: run the pipeline in a continuous job when data should always be current, and let the job own the run mode.
FAQs
Are declarative pipelines and orchestration competing choices?
No. They solve different problems. A declarative pipeline manages relationships between datasets; orchestration manages relationships between tasks. Most production setups use both, with a pipeline running as a task inside a job.
How does a pipeline run inside a job?
It runs as a pipeline task. In a triggered or scheduled job, the task starts a single pipeline update and stops when it completes. In a continuous job, the task keeps the pipeline running.
Do I write If/else logic inside a declarative pipeline?
No. Conditional branches and loops belong in the orchestration layer, as If/else and For each tasks in a job. Inside a pipeline you declare datasets and let the engine resolve execution.
What kinds of tasks can a job orchestrate?
A job can run notebooks, SQL, Python scripts, dashboards, dbt projects, and pipelines, arranged as a directed acyclic graph with schedules, triggers, retries, and monitoring.
Is Lakeflow generally available?
Yes. Lakeflow, including Lakeflow Pipelines and Lakeflow Jobs, is generally available.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.