Skip to main content Skip to footer

A global chemicals major operating 12+ plant sites across North America required a scalable, automated data ingestion platform to unify operational, engineering, and enterprise data into Cognite Data Fusion (CDF). Data was scattered across SAP PM/MM, IP21 time-series historians, GOS, SPF, and engineering document systems — with no repeatable framework for deploying and managing data pipelines across multiple environments.

Infosys built a Databricks-powered data engineering platform using Databricks Asset Bundles (DAB) — an infrastructure-as-code approach that defines 150+ automated scheduled workflows as YAML code, deployed consistently to dev, staging, and prod via Azure DevOps CI/CD pipelines.

Unity Catalog with Delta Lake serves as the governed, three-tier data layer (DEV / PROD) decoupling source extraction from CDF ingestion. Per-site transformation notebooks (tr_* naming convention) use state-based incremental loading to process only changed records — minimizing compute cost and maximizing data freshness.

The result: every pipeline is version-controlled, environment-safe, and reproducible. The platform covers 12+ plant sites with a single unified codebase and delivers trusted operational intelligence to field workers, operations teams, and engineers across the organization.

150+

Automated Databricks workflows deployed as code

12+

Plant sites with active data pipelines

5+

Source system types integrated (SAP, IP21, GOS, SPF, Docs)

2

Environments managed (DEV / PROD) from a single codebase

Key Challenges

  • Siloed engineering, operational, and IT data across systems.
  • Multiple disparate enterprise and operational data sources.
  • Lack of contextualized relationships between business data assets.
  • Limited access to trusted operational information.
  • Difficulty generating insights from disconnected data environments.
  • Need for enterprise-wide operational visibility.
  • Requirement for scalable data foundation supporting AI initiatives.

Ready to experience?

Talk To Experts

Infosys Approach

Adopted Databricks Asset Bundles (DAB) as the infrastructure-as-code framework for all Databricks Workflow definitions.

Defined 150+ Databricks Jobs as YAML files in a resources/ folder, each parameterized with ${var.*} variables for environment switching.

Established dev, staging, and prod targets in databricks.yml — a single YAML controls workspace host, repo path, catalog, and schedule pause state per environment.

Leveraged Unity Catalog with Delta Lake as a three-tier governed data layer (DEV /PROD) decoupling source ingestion from CDF loading.

Implemented state-based incremental loading using watermark tables — only new or changed records processed per run, with a forceFullLoad widget for full reloads.

Standardized per-site, per-data-type transformation notebooks (tr_* naming convention) covering assets, work orders, notifications, functional locations, and time-series.

Automated CI/CD deployment via Azure DevOps: databricks bundle validates on PR, databricks bundle deploys to target on merge.

Provisioned per-job ephemeral single-node clusters (Standard_F8s) — compute spins up on demand and terminates after each run, minimizing idle cost.

Created cross-site Manufacturing Hub aggregation notebooks producing unified Delta tables for enterprise-wide analytics and Power BI dashboards.

The Solution

Databricks Asset Bundles delivering governed, automated, and reproducible data pipelines at scale.

  • Databricks Asset Bundle (DAB) framework: 150+ workflow jobs as YAML code — ${var.*} substitution makes the same definition run identically in dev and prod.
  • Unity Catalog three-tier Delta Lake (DEV /PROD): source data lands in Delta tables; transformation notebooks read from the governed source catalog and write to the output catalog.
  • Per-site transformation notebooks (tr_* convention): one notebook per site × data type — assets, work orders, notifications, functional locations, process tags.
  • State-based incremental loading with watermark tracking: only new/changed records processed per run; forceFullLoad widget enables complete refresh on demand.
  • Azure DevOps CI/CD pipeline: bundle validate on every PR; bundle deploy on merge — no pipeline change reaches production without code review and validation.
  • Ephemeral single-node job clusters (Standard_F8s, Spark 17.3): spin up per run, shut down on completion — no idle compute cost.
  • Azure Key Vault secret scope: service principal credentials fetched at runtime via dbutils.secrets.get — zero hardcoded credentials in any notebook.
  • Manufacturing Hub aggregation layer: cross-site unified Delta tables powering enterprise KPI dashboards in Power BI.
  • Enabled trusted operational intelligence, advanced analytics, and future Generative AI initiatives on the CDF knowledge graph.

Business Outcomes

150+ automated Databricks workflows deployed as version-controlled YAML code — zero manual deployment steps.

Data pipelines running across 12+ plant sites on a single unified Databricks platform.

Environment-safe deployments: CI/CD pipeline validates every change before reaching production.

State-based incremental loading reduced redundant data processing, cutting compute costs and improving data freshness.

New site onboarding accelerated — adding a site requires one YAML workflow definition and one tr_* notebook, no infrastructure changes.

Established a governed, scalable data foundation enabling advanced analytics, AI use cases, and future Generative AI applications on CDF.

Benefits

Scalable, code-driven data pipelines delivering trusted operational data to every plant site.

Infrastructure as code: every pipeline defined, versioned, peer-reviewed, and deployed through Git and Azure DevOps.

Environment consistency: identical codebase in dev and prod — ${var.*} variable substitution eliminates environment drift.

Cost-efficient compute: ephemeral job clusters spin up on demand and terminate after each run — no idle costs.

Operational data freshness: incremental watermark loading ensures near-real-time data in CDF without expensive full reloads.

Scalable site expansion: onboarding a new plant site requires only a new YAML workflow and a tr_* notebook — no infrastructure redesign.

Security by design: Azure Key Vault secret scopes eliminate hardcoded credentials across every notebook.

Foundation for AI: governed Delta Lake + CDF knowledge graph ready for advanced analytics and Generative AI applications.

Request for services

Find out more about how we can help your organization navigate its next. Let us know your areas of interest so that we can serve you better.

All the fields marked with * are required

You must read and agree to the Privacy Statement before submitting
Please fill all required fields

Thank you for connecting with us. We will respond to you shortly.