Microsoft DP-700: Databricks Asset Bundles on Azure

Databricks Asset Bundles have evolved into what Azure Databricks now calls Declarative Automation Bundles, but the engineering goal is the same: define a data or AI project as deployable source rather than as a collection of resources edited manually in the workspace. A bundle can include project files plus resource definitions for jobs, pipelines, and related Databricks assets, with targets for development, staging, and production.

This makes bundles especially useful in mixed Microsoft Fabric engineering estates where Databricks handles part of the data platform. The question is not whether Fabric and Databricks can coexist; it is whether both sides have a repeatable deployment boundary. Declarative Automation Bundles provide that boundary for Databricks resources.

The existing Fabric version-control discussion has a parallel here: source control is valuable because it turns workspace state into something that can be reviewed, tested, and reproduced.

A bundle should represent a deployable project boundary

The bundle configuration is centered on files such as databricks.yml and resource definitions. It can describe jobs, pipelines, dashboards, models, and other supported resources along with the source files they depend on. This is more useful than storing notebooks in Git while leaving production job configuration as an undocumented set of clicks.

A good bundle boundary groups resources that change together. If a Lakeflow Job and a pipeline are released as one product, managing them in the same bundle makes promotion and rollback easier to reason about. If two systems have independent owners and release cycles, forcing them into one bundle creates unnecessary coupling.

The bundle should therefore mirror product ownership, not workspace layout. A workspace is an environment; a bundle is a versioned deployable unit.

Targets separate environment-specific configuration from shared source

Bundles support deployment targets such as dev and prod. A target can change workspace, resource naming, permissions, variables, and other environment-specific settings while keeping the core project definition shared. That reduces the temptation to maintain separate copies of the same notebooks or YAML for each environment.

Environment differences should be deliberate and minimal. Secrets, catalogs, service principals, compute policies, and workspace URLs may vary, but the transformation logic should not quietly diverge. If production needs a different code path, that difference should be represented explicitly in configuration and tests rather than hidden in manual workspace edits.

This is where the Fabric CI/CD mindset aligns well with Databricks. Both platforms benefit from clear promotion boundaries even though the implementation mechanisms differ.

Validate before deploy, and deploy before run

Databricks recommends bundle validation as part of CI/CD before deployment. Validation can catch configuration errors before they reach a workspace. Deployment then synchronizes the declared resources to the target environment, and jobs or pipelines can be run after that deployment succeeds.

This ordering matters because “code exists in Git” does not mean the target workspace reflects that version. A deployment record gives operators a concrete point to associate with a release. If a production failure occurs, the team can identify which source revision and bundle configuration produced the affected resources.

The workspace UI can also deploy bundles, but automated production promotion is more robust when the target workspace and identity are defined in CI/CD. Manual deployment remains useful for development and troubleshooting, not as the only production control.

Run production deployments with durable identities

Production resources should not depend on the personal account of a developer who happened to create them. Databricks recommends service principals for production jobs because employee accounts change, leave organizations, and carry permissions that are often broader than an automated workload needs.

The same principle should apply to bundle deployment. The CI/CD identity needs enough permission to create or update the resources in the target, but it should not become an all-powerful workspace admin by default. Resource permissions and data permissions should be reviewed separately from deployment permission.

Unity Catalog governance provides the data-side context. A bundle can deploy a job, but Unity Catalog still decides whether that job can access the tables, volumes, models, and functions it needs.

Resource ownership should survive source-controlled deployment

One risk in infrastructure-style deployment is assuming that every resource should be recreated freely. Production jobs can have run history, schedules, permissions, and operational integrations that matter. Bundle identity and deployment behavior help Databricks update declared resources without duplicating them when the same bundle is redeployed.

Teams should still understand what deletion means. Removing a resource from a bundle or deleting a bundle can affect managed resources, while source files may remain. Production changes should therefore be reviewed with the same care as application infrastructure changes.

The safest pattern is to make destructive changes explicit. A pull request that removes a job should be visibly different from one that edits a notebook. Reviewers should understand whether the deployment is changing code, runtime configuration, permissions, or the existence of a production resource.

Bundles are strongest when paired with tests and artifact versioning

Databricks CI/CD guidance recommends compiling and testing code before deployment and using versioned artifacts where appropriate. A bundle can reference a library artifact tied to a Git commit or semantic version, which makes the deployed runtime traceable to reviewed source.

Notebook-only workflows still benefit from unit tests around pure transformation logic, schema checks, and smoke tests against a development catalog. Bundle validation confirms configuration shape; it does not prove the transformation is correct.

For data pipelines, deployment tests should also check permissions and expected source availability. A perfectly valid job definition can fail immediately if the target catalog, secret, or connection is missing in production.

Do not use bundles to hide uncontrolled workspace state

A common anti-pattern is to adopt bundles while allowing production resources to keep changing manually. The bundle then becomes one partial description of reality rather than the authoritative deployment path. The next deployment can overwrite manual fixes or fail because workspace state drifted from source.

Production teams should decide which resources are bundle-managed and treat manual edits as emergency exceptions that must be reconciled back to source. The same discipline applies to Fabric deployment pipelines and Git integration: repeatability comes from reducing undocumented state, not merely adding a Git repository beside it.

Declarative Automation Bundles are valuable because they make Databricks releases reviewable. They do not replace architecture, data governance, or testing, but they give those practices a reliable deployment mechanism.

Use the bundle as the handoff between development and operations

The clean operational outcome is that a new engineer can inspect the repository and understand what will be deployed, where it will run, which resources it creates, and how to validate it before production. That is a stronger handoff than a checklist describing which workspace buttons to click.

In a broader Azure data platform, this also gives Fabric and Databricks teams a common engineering language: versioned source, environment targets, tested promotion, durable identities, and observable production resources. The tools differ, but the control objective is the same.

Permissions also deserve environment-specific testing. A developer can successfully validate and deploy a bundle into a personal development workspace while the same bundle fails in production because the service principal lacks permission on a Unity Catalog object, secret scope, connection, or compute policy. A preproduction deployment should verify not only that resources can be created but that the runtime identity can execute the exact data path the job will use.

Bundle variables are useful for keeping environment differences explicit, but they should not become a second programming language full of hidden conditional logic. A small number of clear variables for catalog names, workspace IDs, service endpoints, and policy identifiers is easier to audit than dozens of values that silently alter behavior. If a target needs radically different logic, the project may actually represent two products rather than two environments of one product.

Deployment state should also be observable after promotion. The production handoff should record bundle identity, target, source revision, deployment time, and the resources changed. That makes rollback and incident review faster because operators can connect a job failure to the release that modified it instead of comparing workspace screenshots.

Finally, CI/CD should test deletion and rename scenarios. Resource renames can look like harmless refactoring in source but may create a new production resource while leaving the old one behind, depending on how identity is represented. Reviewers should treat name and key changes as lifecycle changes, verify the deployment plan, and confirm that schedules, permissions, alerts, and downstream references still point to the intended resource.

Repository structure should reinforce the deployment model. Keep bundle configuration close to the code it deploys, document the expected target names, and avoid committing environment secrets. Reviewers should be able to tell which job, pipeline, or dashboard a source change can affect without opening the production workspace. That visibility is one of the main advantages of treating Databricks resources as source-controlled configuration.

As the bundle estate grows, teams should standardize templates for recurring patterns such as naming, tags, permissions, cluster policy references, and observability hooks. Templates reduce accidental variation without forcing every project into one monolith. The standard should make good defaults easy while still letting a project override the parts that genuinely differ.

Leave a Reply

How It Works

img
Step 1. Choose Exam
on ExamLabs
Download IT Exams Questions & Answers
img
Step 2. Open Exam with
Avanset Exam Simulator
Press here to download VCE Exam Simulator that simulates real exam environment
img
Step 3. Study
& Pass
IT Exams Anywhere, Anytime!