Case Study
Chemicals Major Automates 150+ Data Pipelines with Databricks
- AI & Automation
- Data Integration
A global chemicals major operating 12+ plant sites across North America required a scalable, automated data ingestion platform to unify operational, engineering, and enterprise data into Cognite Data Fusion (CDF). Data was scattered across SAP PM/MM, IP21 time-series historians, GOS, SPF, and engineering document systems — with no repeatable framework for deploying and managing data pipelines across multiple environments.
Infosys built a Databricks-powered data engineering platform using Databricks Asset Bundles (DAB) — an infrastructure-as-code approach that defines 150+ automated scheduled workflows as YAML code, deployed consistently to dev, staging, and prod via Azure DevOps CI/CD pipelines.
Unity Catalog with Delta Lake serves as the governed, three-tier data layer (DEV / PROD) decoupling source extraction from CDF ingestion. Per-site transformation notebooks (tr_* naming convention) use state-based incremental loading to process only changed records — minimizing compute cost and maximizing data freshness.
The result: every pipeline is version-controlled, environment-safe, and reproducible. The platform covers 12+ plant sites with a single unified codebase and delivers trusted operational intelligence to field workers, operations teams, and engineers across the organization.
Automated Databricks workflows deployed as code
Plant sites with active data pipelines
Source system types integrated (SAP, IP21, GOS, SPF, Docs)
Environments managed (DEV / PROD) from a single codebase
Ready to experience?
Talk To ExpertsAdopted Databricks Asset Bundles (DAB) as the infrastructure-as-code framework for all Databricks Workflow definitions.
Defined 150+ Databricks Jobs as YAML files in a resources/ folder, each parameterized with ${var.*} variables for environment switching.
Established dev, staging, and prod targets in databricks.yml — a single YAML controls workspace host, repo path, catalog, and schedule pause state per environment.
Leveraged Unity Catalog with Delta Lake as a three-tier governed data layer (DEV /PROD) decoupling source ingestion from CDF loading.
Implemented state-based incremental loading using watermark tables — only new or changed records processed per run, with a forceFullLoad widget for full reloads.
Standardized per-site, per-data-type transformation notebooks (tr_* naming convention) covering assets, work orders, notifications, functional locations, and time-series.
Automated CI/CD deployment via Azure DevOps: databricks bundle validates on PR, databricks bundle deploys to target on merge.
Provisioned per-job ephemeral single-node clusters (Standard_F8s) — compute spins up on demand and terminates after each run, minimizing idle cost.
Created cross-site Manufacturing Hub aggregation notebooks producing unified Delta tables for enterprise-wide analytics and Power BI dashboards.
Databricks Asset Bundles delivering governed, automated, and reproducible data pipelines at scale.
150+ automated Databricks workflows deployed as version-controlled YAML code — zero manual deployment steps.
Data pipelines running across 12+ plant sites on a single unified Databricks platform.
Environment-safe deployments: CI/CD pipeline validates every change before reaching production.
State-based incremental loading reduced redundant data processing, cutting compute costs and improving data freshness.
New site onboarding accelerated — adding a site requires one YAML workflow definition and one tr_* notebook, no infrastructure changes.
Established a governed, scalable data foundation enabling advanced analytics, AI use cases, and future Generative AI applications on CDF.
Scalable, code-driven data pipelines delivering trusted operational data to every plant site.
Infrastructure as code: every pipeline defined, versioned, peer-reviewed, and deployed through Git and Azure DevOps.
Environment consistency: identical codebase in dev and prod — ${var.*} variable substitution eliminates environment drift.
Cost-efficient compute: ephemeral job clusters spin up on demand and terminate after each run — no idle costs.
Operational data freshness: incremental watermark loading ensures near-real-time data in CDF without expensive full reloads.
Scalable site expansion: onboarding a new plant site requires only a new YAML workflow and a tr_* notebook — no infrastructure redesign.
Security by design: Azure Key Vault secret scopes eliminate hardcoded credentials across every notebook.
Foundation for AI: governed Delta Lake + CDF knowledge graph ready for advanced analytics and Generative AI applications.
Find out more about how we can help your organization navigate its next. Let us know your areas of interest so that we can serve you better.