Modern Data

Data Platform Foundation

A starting data platform, set up as code on your cloud.

A reference architecture and infrastructure-as-code templates for a lakehouse or cloud data warehouse. It covers ingestion, storage layers, access policies, and deployment pipelines. We fit it to your cloud and the tools you already use, so your team can spend its time building data products.

GIST stepIterateShip
Download the PDF
What you get

A working data platform your team can run and extend.

Free download · PDF · 10 pages

Get the Data Platform Foundation overview

The problem it solves, what is included, how it works, the technical components, and how we adapt it with you.

The problem

Platform setup eats the first months of a data program.

Networking, storage, identity, access policies, and deployment pipelines have to be right before anyone builds a data product. Teams rebuild them on every program, and security settings get decided late.

What this accelerator does

The Data Platform Foundation gives you a governed lakehouse as code, fitted to your cloud, so your team starts on data products.

What's included

A governed lakehouse, set up as code.

Four parts, each adapted to your data, platforms, and controls. What we adapt for you is yours to keep.

01

Reference architecture

For Azure, AWS, and Google Cloud, with Microsoft Fabric, Snowflake, and BigQuery variants, plus an access model.

02

Infrastructure as code

Terraform modules for networking, storage, Key Vault, Databricks, Unity Catalog, and monitoring, per environment.

03

Ingestion and orchestration

Templates for files, incremental database loads, and REST APIs, then typed and tested silver and gold layers.

04

Controls and pipelines

Access grants per layer, quality rules that quarantine bad rows, guardrail checks, and CI/CD with production approval.

How it works

What gets built on Azure.

  1. Private network

    A virtual network with private endpoints, delegated Databricks subnets, and private DNS zones.

  2. Governed lake

    ADLS Gen2 with a container per layer: landing, bronze, silver, gold. Shared keys and public access off.

  3. Secrets

    Key Vault with role-based access only, network deny, and purge protection in production.

  4. Compute

    A Premium Databricks workspace with VNet injection and no public IPs; private front end in production.

  5. Access as code

    Unity Catalog grants: admins all, engineers read-write, analysts read silver and gold.

  6. Monitoring

    Storage, Key Vault, and Databricks audit logs sent to Log Analytics.

Technical detail

Under the hood.

Vendor-neutral Python and configuration, Azure first, with tests included from the start.

Terraform modules
Network, storage, Key Vault, Databricks, Unity Catalog, and monitoring
Environments
Separate dev and prod settings, each with its own remote state
Databricks Asset Bundle
Ingestion and transformation jobs, validated and deployed per target
Source templates
Auto Loader for files, JDBC with a watermark, and paginated REST APIs
Quality rules
Rules in YAML, a quarantine table, and a threshold that stops bad loads
Guardrails and CI/CD
Security checks in every pull request; OIDC login and production approval
Proof in the package

Security settings are checked before anything is deployed.

Guardrail checks run in every pull request with no cloud access needed, so problems surface in review instead of after go-live.

What the checks cover
  • Wiring between network, storage, and compute modules
  • Environment settings files, complete and consistent
  • Security settings from the reference design
  • Quality rules applied to every source load
Result

A change that weakens a security setting fails in the pull request, before it reaches a subscription. Production deployments wait for an approval.

How we run it with you

Adapted in the first cycles, handed over at the end.

  1. 01

    Names and network

    Naming, region, and an address space that fits your hub-and-spoke network and DNS.

  2. 02

    Groups and access

    Your Entra ID groups for admins, engineers, and analysts mapped to catalog grants.

  3. 03

    Sources

    A template per source, credentials in Key Vault, and quality rules agreed with each owner.

  4. 04

    Policies and clouds

    Your Azure Policy assignments added. For AWS or Google Cloud, we write those modules.

You keep the Terraform, the Databricks bundle, the quality rules, and the pipelines, in your repository and run by your team.

Where it fits

Platforms, related accelerators, and limits.

We state the limits up front, and we recommend tools based on fit. We do not resell platforms.

Works with

  • Azure first: ADLS Gen2, Azure Databricks, Unity Catalog, Key Vault
  • Reference designs for AWS and Google Cloud
  • Fabric, Snowflake, and BigQuery variants

Pairs with

  • Runs the pipelines, master data, and metrics from the other accelerators
  • MLOps Templates train and monitor models on it
  • Quality rules become part of each source's data contract

Assumptions and limits

  • Static checks pass today; validating against your subscription is the first step
  • One environment per state file; account-level resources are set up once
  • Databricks, storage, and logging have running costs we size with you first

Take the overview with you.

Get the 10-page PDF to share with your team, or tell us the decision you want to improve and we will tell you whether the Data Platform Foundation fits.