Reference architecture
For Azure, AWS, and Google Cloud, with Microsoft Fabric, Snowflake, and BigQuery variants, plus an access model.
A starting data platform, set up as code on your cloud.
A reference architecture and infrastructure-as-code templates for a lakehouse or cloud data warehouse. It covers ingestion, storage layers, access policies, and deployment pipelines. We fit it to your cloud and the tools you already use, so your team can spend its time building data products.
Download the PDFA working data platform your team can run and extend.
The problem it solves, what is included, how it works, the technical components, and how we adapt it with you.
Your download has started.
We also emailed the link to . It stays valid for 7 days.
Download didn't start? Get the PDF
Want to see how it would fit your data? Talk to us.
Networking, storage, identity, access policies, and deployment pipelines have to be right before anyone builds a data product. Teams rebuild them on every program, and security settings get decided late.
The Data Platform Foundation gives you a governed lakehouse as code, fitted to your cloud, so your team starts on data products.
Four parts, each adapted to your data, platforms, and controls. What we adapt for you is yours to keep.
For Azure, AWS, and Google Cloud, with Microsoft Fabric, Snowflake, and BigQuery variants, plus an access model.
Terraform modules for networking, storage, Key Vault, Databricks, Unity Catalog, and monitoring, per environment.
Templates for files, incremental database loads, and REST APIs, then typed and tested silver and gold layers.
Access grants per layer, quality rules that quarantine bad rows, guardrail checks, and CI/CD with production approval.
A virtual network with private endpoints, delegated Databricks subnets, and private DNS zones.
ADLS Gen2 with a container per layer: landing, bronze, silver, gold. Shared keys and public access off.
Key Vault with role-based access only, network deny, and purge protection in production.
A Premium Databricks workspace with VNet injection and no public IPs; private front end in production.
Unity Catalog grants: admins all, engineers read-write, analysts read silver and gold.
Storage, Key Vault, and Databricks audit logs sent to Log Analytics.
Vendor-neutral Python and configuration, Azure first, with tests included from the start.
Guardrail checks run in every pull request with no cloud access needed, so problems surface in review instead of after go-live.
A change that weakens a security setting fails in the pull request, before it reaches a subscription. Production deployments wait for an approval.
Naming, region, and an address space that fits your hub-and-spoke network and DNS.
Your Entra ID groups for admins, engineers, and analysts mapped to catalog grants.
A template per source, credentials in Key Vault, and quality rules agreed with each owner.
Your Azure Policy assignments added. For AWS or Google Cloud, we write those modules.
You keep the Terraform, the Databricks bundle, the quality rules, and the pipelines, in your repository and run by your team.
We state the limits up front, and we recommend tools based on fit. We do not resell platforms.
Get the 10-page PDF to share with your team, or tell us the decision you want to improve and we will tell you whether the Data Platform Foundation fits.