Tutorial: How to Orchestrate Databricks Jobs with Airflow

Databricks is a popular unified data and analytics platform built around Apache Spark that provides users with fully managed Apache Spark clusters and interactive workspaces.

The open source Airflow Databricks provider provides full observability and control from Airflow so you can manage Databricks from one place, including enabling you to orchestrate your Databricks notebooks from Airflow and execute them as Databricks jobs.

Why use Airflow with Databricks

Many data teams leverage Databricks’ optimized Spark engine to run heavy workloads like machine learning models, data transformations, and data analysis. While Databricks offers some orchestration with Databricks Workflows, they are limited in functionality and do not integrate with the rest of your data stack. Using a tool-agnostic orchestrator like Airflow gives you several advantages, like the ability to:

Use CI/CD to manage your workflow deployment. Airflow Dags are Python code, and can be integrated with a variety of CI/CD tools and tested.
Use task groups within Databricks jobs, enabling you to collapse and expand parts of larger Databricks jobs visually.
Leverage Airflow assets to trigger Databricks jobs from tasks in other Dags in your Airflow environment or using the Airflow REST API Create asset event endpoint, allowing for a data-driven architecture.
Use familiar Airflow code as your interface to orchestrate Databricks notebooks as jobs.
Inject parameters into your Databricks job at the job-level. These parameters can be dynamic and retrieved at runtime from other Airflow tasks.
Directly jump from a task in your Airflow Dag to the corresponding Databricks job in the Databricks UI using an operator extra link.

Time to complete

This step-by-step tutorial takes approximately 30 minutes to complete. After completting this tutorial, you will have a working Airflow Dag that orchestrates Databricks notebooks as a Databricks Workflow.

Assumed knowledge

To get the most out of this tutorial, make sure you have an understanding of:

The basics of Databricks. See Getting started with Databricks.
Airflow fundamentals, such as writing Dags and defining tasks. See Get started with Apache Airflow.
Airflow operators. See Operators 101.
Airflow connections. See Managing your Connections in Apache Airflow.

Prerequisites

The Astro CLI.
Access to a Databricks workspace. See Databricks’ documentation for instructions. You can use any workspace that has access to the Databricks Workflows feature. You need a user account with permissions to create notebooks and Databricks jobs. You can use any underlying cloud service, and a 14-day free trial is available.

Step 1: Configure your Astro project

Create a new Astro project:

1 $ mkdir astro-databricks-tutorial && cd astro-databricks-tutorial
2 $ astro dev init

Add the Airflow Databricks provider package to your requirements.txt file.

apache-airflow-providers-databricks==7.8.0

Step 2: Create Databricks Notebooks

You can orchestrate any Databricks notebooks in a Databricks job using the Airflow Databricks provider. If you don’t have Databricks notebooks ready, follow these steps to create two notebooks:

Create an empty notebook in your Databricks workspace called notebook1.

Copy and paste the following code into the first cell of the notebook1 notebook.

1 print("Hello")

Create a second empty notebook in your Databricks workspace called notebook2.

Copy and paste the following code into the first cell of the notebook2 notebook.

1 print("World")

Step 3: Configure the Databricks connection

Start Airflow by running astro dev start.
In the Airflow UI, go to Admin > Connections and click +.
Create a new connection named databricks_conn. Select the connection type Databricks and enter the following information:
- Connection ID: databricks_conn.
- Connection Type: Databricks.
- Host: Your Databricks host address (format: https://dbc-1234cb56-d7c8.cloud.databricks.com/).
- Password: Your Databricks personal access token.
Alternatively, you can create an OAuth connection to your Databricks workspace by providing the Host, Service Principal Client ID as Login, Service Principal Client Secret as Password and set service_principal_oauth to True in the Extra field.

Astro customers can use the Astro Environment Manager to create a connection to Databricks, stored in the Astro-managed secrets backend. This connection can be shared across multiple Deployments in a Workspace.

Step 4: Create your Dag

In your dags folder, create a file called my_simple_databricks_dag.py.

Copy and paste the following Dag code into the file. Replace<your-databricks-login-email> variable with your Databricks login email. If you already had Databricks notebooks and did not create new ones in Step 2, adjust the notebook_path parameters in the two DatabricksNotebookOperators to point to the existing notebooks. Adjust the job_cluster_spec to match your available cloud resources.

1 """
2 ### Run notebooks in databricks as a Databricks Workflow using the Airflow Databricks provider
3 
4 This Dag runs two Databricks notebooks as a Databricks workflow.
5 """
6 
7 from airflow.sdk import dag, chain
8 from airflow.providers.databricks.operators.databricks import DatabricksNotebookOperator
9 from airflow.providers.databricks.operators.databricks_workflow import (
10     DatabricksWorkflowTaskGroup,
11 )
12 from pendulum import datetime
13 
14 DATABRICKS_LOGIN_EMAIL = "<your-databricks-login-email>"
15 DATABRICKS_NOTEBOOK_NAME_1 = "notebook1"
16 DATABRICKS_NOTEBOOK_NAME_2 = "notebook2"
17 DATABRICKS_NOTEBOOK_PATH_1 = (
18     f"/Users/{DATABRICKS_LOGIN_EMAIL}/{DATABRICKS_NOTEBOOK_NAME_1}"
19 )
20 DATABRICKS_NOTEBOOK_PATH_2 = (
21     f"/Users/{DATABRICKS_LOGIN_EMAIL}/{DATABRICKS_NOTEBOOK_NAME_2}"
22 )
23 DATABRICKS_JOB_CLUSTER_KEY = "tutorial-cluster"
24 DATABRICKS_CONN_ID = "databricks_conn"
25 
26 # adjust if necessary for example to align the spark version with your Notebooks
27 job_cluster_spec = [
28     {
29         "job_cluster_key": DATABRICKS_JOB_CLUSTER_KEY,
30         "new_cluster": {
31             "cluster_name": "",
32             "spark_version": "15.4.x-scala2.12",
33             "azure_attributes": {
34                 "first_on_demand": 1,
35                 "availability": "SPOT_WITH_FALLBACK_AZURE",
36                 "spot_bid_max_price": -1,
37             },
38             "node_type_id": "Standard_DS3_v2",
39             "spark_env_vars": {"PYSPARK_PYTHON": "/databricks/python3/bin/python3"},
40             "enable_elastic_disk": False,
41             "data_security_mode": "LEGACY_SINGLE_USER_STANDARD",
42             "runtime_engine": "STANDARD",
43             "num_workers": 1,
44         },
45     }
46 ]
47 
48 
49 @dag
50 def my_simple_databricks_dag():
51     task_group = DatabricksWorkflowTaskGroup(
52         group_id="databricks_workflow",
53         databricks_conn_id=DATABRICKS_CONN_ID,
54         job_clusters=job_cluster_spec,
55     )
56 
57     with task_group:
58         notebook_1 = DatabricksNotebookOperator(
59             task_id="notebook1",
60             databricks_conn_id=DATABRICKS_CONN_ID,
61             notebook_path=DATABRICKS_NOTEBOOK_PATH_1,
62             source="WORKSPACE",
63             job_cluster_key=DATABRICKS_JOB_CLUSTER_KEY,
64         )
65         notebook_2 = DatabricksNotebookOperator(
66             task_id="notebook2",
67             databricks_conn_id=DATABRICKS_CONN_ID,
68             notebook_path=DATABRICKS_NOTEBOOK_PATH_2,
69             source="WORKSPACE",
70             job_cluster_key=DATABRICKS_JOB_CLUSTER_KEY,
71         )
72         chain(notebook_1, notebook_2)
73 
74 
75 my_simple_databricks_dag()

This Dag uses the Airflow Databricks provider to create a Databricks job that runs two notebooks. The databricks_workflow task group, created using the DatabricksWorkflowTaskGroup class, automatically creates a Databricks job that executes the Databricks notebooks you specified in the individual DatabricksNotebookOperators. One of the biggest benefits of this setup is the use of a Databricks job cluster, allowing you to significantly reduce your Databricks cost. The task group contains three tasks:

The launch task, which the task group automatically generates, provisions a Databricks job_cluster with the spec defined as job_cluster_spec and creates the Databricks job from the tasks within the task group.
The notebook1 task runs the notebook1 notebook in this cluster as the first part of the Databricks job.
The notebook2 task runs the notebook2 notebook as the second part of the Databricks job.

Run the Dag manually by clicking the play button and view the Dag in the graph tab. In case the task group appears collapsed, click it in order to expand and see all tasks.
View the completed Databricks job in the Databricks UI.

Step 5 (optional): Add a task to run SQL

You can run any SQL query in Databricks using the DatabricksSqlOperator from the Airflow Databricks provider. In your Dag, outside of the databricks_workflow task group, add the following task. Replace the placeholder values with your own values.

1 from airflow.providers.databricks.operators.databricks_sql import DatabricksSqlOperator
2 
3 run_sql = DatabricksSqlOperator(
4     task_id="run_sql",
5     databricks_conn_id=DBX_CONN_ID,
6     http_path=f"/sql/1.0/warehouses/{DATABRICKS_SQL_WAREHOUSE_ID}",
7     catalog="<your-catalog-name>",
8     schema="<your-schema-name>",
9     sql="<your-sql-query>",
10     parameters={"<your-parameter-name>": "<your-parameter-value>"}
11 )

Alternatively, you can also use the DatabricksHook directly in any @task decorated function or PythonOperator in your Dag.

1     @task
2     def run_sql():
3         from airflow.providers.databricks.hooks.databricks import DatabricksHook
4 
5         hook = DatabricksHook(DBX_CONN_ID)
6 
7         re = hook.post_sql_statement(
8             json={
9                 "warehouse_id": DATABRICKS_SQL_WAREHOUSE_ID,
10                 "catalog": "<your-catalog-name>",
11                 "schema": "<your-schema-name>",
12                 "statement": "<your-sql-query>",
13                 "parameters": [
14                     {"name": "<your-parameter-name>", "value": "<your-parameter-value>", "type": "<your-parameter-type>"},
15                 ],
16             }
17         )

How it works

This section explains Airflow Databricks provider functionality in more depth. You can learn more about the Airflow Databricks provider, including more information about other available operators, in the provider documentation.

Parameters

The DatabricksWorkflowTaskGroup provides configuration options via several parameters:

job_clusters: the job clusters parameters for this job to use. You can provide the full job_cluster_spec as shown in the tutorial Dag.

notebook_params: a dictionary of parameters to make available to all notebook tasks in a job. This operator is templatable, see below for a code example:

1 dbx_workflow_task_group = DatabricksWorkflowTaskGroup(
2     group_id="databricks_workflow",
3     databricks_conn_id=_DBX_CONN_ID,
4     job_clusters=job_cluster_spec,
5     notebook_params={
6         "my_date": "{{ ds }}"
7     },
8 )

To retrieve this parameter inside your Databricks notebook add the following code to a Databricks notebook cell:

1 dbutils.widgets.text("my_date", "my_default_value", "Description")
2 my_date = dbutils.widgets.get("my_date")

notebook_packages: a list of dictionaries defining Python packages to install in all notebook tasks in a job.
extra_job_params: a dictionary with properties to override the default Databricks job definitions.

You also have the ability to specify parameters at the task level in the DatabricksNotebookOperator:

notebook_params: a dictionary of parameters to make available to the notebook.
notebook_packages: a list of dictionaries defining Python packages to install in the notebook.

Note that you cannot specify the same packages in both the notebook_packages parameter of a DatabricksWorkflowTaskGroup and the notebook_packages parameter of a task using the DatabricksNotebookOperator in that same task group. Duplicate entries in this parameter cause an error in Databricks.

1	$ mkdir astro-databricks-tutorial && cd astro-databricks-tutorial
2	$ astro dev init

1	"""
2	### Run notebooks in databricks as a Databricks Workflow using the Airflow Databricks provider
3
4	This Dag runs two Databricks notebooks as a Databricks workflow.
5	"""
6
7	from airflow.sdk import dag, chain
8	from airflow.providers.databricks.operators.databricks import DatabricksNotebookOperator
9	from airflow.providers.databricks.operators.databricks_workflow import (
10	DatabricksWorkflowTaskGroup,
11	)
12	from pendulum import datetime
13
14	DATABRICKS_LOGIN_EMAIL = "<your-databricks-login-email>"
15	DATABRICKS_NOTEBOOK_NAME_1 = "notebook1"
16	DATABRICKS_NOTEBOOK_NAME_2 = "notebook2"
17	DATABRICKS_NOTEBOOK_PATH_1 = (
18	f"/Users/{DATABRICKS_LOGIN_EMAIL}/{DATABRICKS_NOTEBOOK_NAME_1}"
19	)
20	DATABRICKS_NOTEBOOK_PATH_2 = (
21	f"/Users/{DATABRICKS_LOGIN_EMAIL}/{DATABRICKS_NOTEBOOK_NAME_2}"
22	)
23	DATABRICKS_JOB_CLUSTER_KEY = "tutorial-cluster"
24	DATABRICKS_CONN_ID = "databricks_conn"
25
26	# adjust if necessary for example to align the spark version with your Notebooks
27	job_cluster_spec = [
28	{
29	"job_cluster_key": DATABRICKS_JOB_CLUSTER_KEY,
30	"new_cluster": {
31	"cluster_name": "",
32	"spark_version": "15.4.x-scala2.12",
33	"azure_attributes": {
34	"first_on_demand": 1,
35	"availability": "SPOT_WITH_FALLBACK_AZURE",
36	"spot_bid_max_price": -1,
37	},
38	"node_type_id": "Standard_DS3_v2",
39	"spark_env_vars": {"PYSPARK_PYTHON": "/databricks/python3/bin/python3"},
40	"enable_elastic_disk": False,
41	"data_security_mode": "LEGACY_SINGLE_USER_STANDARD",
42	"runtime_engine": "STANDARD",
43	"num_workers": 1,
44	},
45	}
46	]
47
48
49	@dag
50	def my_simple_databricks_dag():
51	task_group = DatabricksWorkflowTaskGroup(
52	group_id="databricks_workflow",
53	databricks_conn_id=DATABRICKS_CONN_ID,
54	job_clusters=job_cluster_spec,
55	)
56
57	with task_group:
58	notebook_1 = DatabricksNotebookOperator(
59	task_id="notebook1",
60	databricks_conn_id=DATABRICKS_CONN_ID,
61	notebook_path=DATABRICKS_NOTEBOOK_PATH_1,
62	source="WORKSPACE",
63	job_cluster_key=DATABRICKS_JOB_CLUSTER_KEY,
64	)
65	notebook_2 = DatabricksNotebookOperator(
66	task_id="notebook2",
67	databricks_conn_id=DATABRICKS_CONN_ID,
68	notebook_path=DATABRICKS_NOTEBOOK_PATH_2,
69	source="WORKSPACE",
70	job_cluster_key=DATABRICKS_JOB_CLUSTER_KEY,
71	)
72	chain(notebook_1, notebook_2)
73
74
75	my_simple_databricks_dag()

1	from airflow.providers.databricks.operators.databricks_sql import DatabricksSqlOperator
2
3	run_sql = DatabricksSqlOperator(
4	task_id="run_sql",
5	databricks_conn_id=DBX_CONN_ID,
6	http_path=f"/sql/1.0/warehouses/{DATABRICKS_SQL_WAREHOUSE_ID}",
7	catalog="<your-catalog-name>",
8	schema="<your-schema-name>",
9	sql="<your-sql-query>",
10	parameters={"<your-parameter-name>": "<your-parameter-value>"}
11	)

1	@task
2	def run_sql():
3	from airflow.providers.databricks.hooks.databricks import DatabricksHook
4
5	hook = DatabricksHook(DBX_CONN_ID)
6
7	re = hook.post_sql_statement(
8	json={
9	"warehouse_id": DATABRICKS_SQL_WAREHOUSE_ID,
10	"catalog": "<your-catalog-name>",
11	"schema": "<your-schema-name>",
12	"statement": "<your-sql-query>",
13	"parameters": [
14	{"name": "<your-parameter-name>", "value": "<your-parameter-value>", "type": "<your-parameter-type>"},
15	],
16	}
17	)

1	dbx_workflow_task_group = DatabricksWorkflowTaskGroup(
2	group_id="databricks_workflow",
3	databricks_conn_id=_DBX_CONN_ID,
4	job_clusters=job_cluster_spec,
5	notebook_params={
6	"my_date": "{{ ds }}"
7	},
8	)

1	dbutils.widgets.text("my_date", "my_default_value", "Description")
2	my_date = dbutils.widgets.get("my_date")