Sathus AI 2.0 is now generally available — evaluation harnesses and guardrails included. Explore
The comprehensive enterprise architecture blueprint for designing, deploying, and operating an Azure Modern Data Platform. Covers Azure Data Lake Storage Gen2 (ADLS Gen2) hierarchical namespaces, Azure Databricks Delta Lake, Microsoft Fabric OneLake integration, and Microsoft Purview data governance.
An Azure Modern Data Platform is an enterprise-scale analytical foundation built on Azure Data Lake Storage Gen2 (ADLS Gen2) hierarchical namespaces, governed universally by Microsoft Purview, and powered by Azure Databricks and Microsoft Fabric compute engines. Raw data from hybrid on-premises and SaaS sources is ingested via Azure Data Factory (ADF) into Bronze ADLS containers, while Azure Databricks handles distributed Silver cleaning and Gold dimensional modeling using Delta Lake. Analytical reporting is delivered through Microsoft Fabric and Power BI Direct Lake mode, which queries Gold Delta tables directly from object storage without semantic model refresh latency or data duplication, all secured under Microsoft Entra ID and private endpoints.
The comprehensive enterprise architecture blueprint for designing, deploying, and operating an Azure Modern Data Platform. Covers Azure Data Lake Storage Gen2 (ADLS Gen2) hierarchical namespaces, Azure Databricks Delta Lake, Microsoft Fabric OneLake integration, and Microsoft Purview data governance.
A: Standard blob storage does not have true directories; a directory is merely a prefix in the object key. Renaming a directory requires copying every single file within it (O(N) complexity). With HNS enabled, directories are true file system objects, enabling atomic, O(1) directory renames—which is essential for Spark commit operations and avoiding partial writes.
A: Import mode delivers top-tier sub-second query performance by loading data into the Power BI Analysis Services cache, but requires scheduled refresh cycles and duplicates data. Direct Lake mode reads directly from Delta Parquet files in ADLS Gen2 / OneLake at in-memory speeds without data duplication and without scheduled refresh delays.
A: Microsoft Purview connects directly to Databricks Unity Catalog via managed metadata connectors. It automatically ingests table schemas, tags, descriptions, and column-level lineage into the enterprise Purview Data Map, ensuring centralized governance across both Azure and multi-cloud services.
Azure Data Lake Storage Gen2 combines the cost-efficiency of Blob Storage with high-performance file system semantics. Storage design standards: • Hierarchical Namespace (HNS): Must be enabled at storage account creation. HNS enables atomic folder renames, eliminating the slow multi-object copy overhead of standard object storage during Spark atomic commits. • Three-Container Medallion Topology: Segregate bronze (raw append-only), silver (cleansed & conformed), and gold (business aggregates) into separate containers with distinct access policies. • Dual-Layer Security: Combine Azure Role-Based Access Control (RBAC) via Microsoft Entra ID with POSIX-compliant Access Control Lists (ACLs) for granular folder-level permissions.
A modern Azure data architecture balances orchestration and compute: • Azure Data Factory (ADF): Utilized for hybrid connectivity (Self-Hosted Integration Runtime for on-premise relational databases), SAP ingestion, and high-level pipeline orchestration across cloud services. • Azure Databricks Workflows: Utilized for heavy transformations, feature engineering, and PySpark Delta Lake processing where deep Spark execution graph tuning is required. • Event-Driven Triggering: Ingest events triggered via Azure Event Hubs or Azure Storage Event Grid for real-time micro-batch processing.
Enterprises frequently navigate the positioning of Microsoft Fabric alongside Azure Databricks: • Azure Databricks: Best suited for complex data engineering, machine learning lifecycle (MLflow), custom distributed packages, and multi-cloud data strategies. • Microsoft Fabric (OneLake & Direct Lake): Best suited for self-service business analytics, departmental citizen data analysts, and executive reporting. • Interoperability: OneLake Shortcuts allow Fabric and Power BI to query Azure Databricks Delta tables in ADLS Gen2 without copying or moving data.
Unified data governance across Azure: • Automated Scanning: Purview automated scanners catalog ADLS Gen2 containers, Azure SQL databases, and Databricks Unity Catalog assets. • Classification Rules: Automatically detects credit card numbers, SSNs, HIPAA medical codes, and GDPR personal identifiers. • End-to-End Lineage: Visualizes the transformation path from source database -> ADF Copy -> Databricks Delta table -> Power BI semantic model.
Production Azure Modern Data Platform spanning ADLS Gen2 storage, Databricks compute, ADF orchestration, and Purview governance.
Hierarchical namespace storage containers secured via Private Endpoints and Entra ID RBAC.
Scalable Spark clusters for Silver/Gold Delta Lake transformations and Unity Catalog governance.
Automated data lineage via Purview and sub-second BI reporting via Power BI Direct Lake.
Pinpoint root cause failure modes and match observed metrics to actionable remediation.
| UI Tab / Tool | Observed Metric / Signal | Underlying Failure Mode | Actionable Remediation |
|---|---|---|---|
| Azure Storage / Diagnostic Logs | HTTP 403 AuthorizationPermissionMismatch when Databricks writes to ADLS Gen2 | Service Principal or Managed Identity lacks the "Storage Blob Data Contributor" role on the target container. | Assign "Storage Blob Data Contributor" role in IAM and verify POSIX execute (x) permissions on parent directories. |
| Azure Data Factory Monitor | Copy Activity throughput drops below 1 MB/s during hybrid on-premises ingestion | Self-Hosted Integration Runtime (SHIR) machine CPU/memory bottleneck or network throttling on ExpressRoute/VPN. | Scale up SHIR VM instance size, increase concurrent copy tasks, and adjust Data Integration Units (DIU). |
| Azure Databricks Workspace | Cluster creation fails with "VNet injection IP exhaustion / SubnetIsFull" | The private and public subnets allocated for Databricks workers ran out of available private IP addresses. | Size worker subnets with at least a /24 (256 IPs) or /23 (512 IPs) CIDR block during initial VNet configuration. |
| Power BI Service | Direct Lake dataset falls back to DirectQuery mode during report rendering | Underlying Delta table contains unsupported data types (e.g. binary/complex structs) or exceeds SKU memory limit. | Flatten complex structs in the Gold layer and verify Delta table file size fits within Fabric capacity memory. |
param location string = resourceGroup().location
param storageAccountName string = 'sathuslakehousegold'
param vnetName string = 'vnet-data-platform'
param subnetName string = 'snet-private-endpoints'
// 1. Create ADLS Gen2 Storage Account with Hierarchical Namespace
resource storageAccount 'Microsoft.Storage/storageAccounts@2023-01-01' = {
name: storageAccountName
location: location
sku: {
name: 'Standard_ZRS'
}
kind: 'StorageV2'
properties: {
isHnsEnabled: true // Enable Hierarchical Namespace (ADLS Gen2)
minimumTlsVersion: 'TLS1_2'
supportsHttpsTrafficOnly: true
publicNetworkAccess: 'Disabled' // Enforce Private Endpoint Access
networkAcls: {
bypass: 'AzureServices'
defaultAction: 'Deny'
}
}
}
// 2. Medallion Storage Containers
resource bronzeContainer 'Microsoft.Storage/storageAccounts/blobServices/containers@2023-01-01' = {
name: '${storageAccount.name}/default/bronze'
}
resource silverContainer 'Microsoft.Storage/storageAccounts/blobServices/containers@2023-01-01' = {
name: '${storageAccount.name}/default/silver'
}
resource goldContainer 'Microsoft.Storage/storageAccounts/blobServices/containers@2023-01-01' = {
name: '${storageAccount.name}/default/gold'
}On-premises SQL Server data warehouse using SQL Server Integration Services (SSIS) and Analysis Services (SSAS) cubes. Storage limits caused nightly ETL runs to balloon to 14 hours, frequently bleeding into business hours.
Decoupled Azure Modern Data Platform on ADLS Gen2, orchestrated by ADF, transformed via Azure Databricks Delta Lake, and served via Power BI Direct Lake mode with sub-second executive dashboard responsiveness.
Reference architecture modeled on enterprise migrations from on-premises Microsoft SQL Server / SSIS to Azure Databricks and ADLS Gen2. Batch window reduction and sub-second Power BI responsiveness reflect typical distributed Spark scaling and Direct Lake in-memory Parquet reading.
| Compute Service | Best Suited For | Primary Strength | Data Access Method | Skillset Required |
|---|---|---|---|---|
| Azure Databricks | Heavy data engineering, ML pipelines | Photon C++ vectorized engine & Spark | Direct Delta Lake in ADLS Gen2 | Python, Scala, Spark SQL |
| Microsoft Fabric | Unified analytics, departmental BI | OneLake Shortcuts & SaaS simplicity | Direct Lake mode on Parquet | SQL, Power BI, DAX |
| Azure Data Factory | Hybrid data movement, SAP, orchestration | 100+ native connectors & SHIR | Managed pipeline copy activities | Low-Code, JSON pipeline spec |
| Azure Synapse (Legacy) | Dedicated SQL DW, classic T-SQL | Massive Parallel Processing (MPP) | PolyBase / External tables | T-SQL, Transact-SQL |
Standard blob storage does not have true directories; a directory is merely a prefix in the object key. Renaming a directory requires copying every single file within it (O(N) complexity). With HNS enabled, directories are true file system objects, enabling atomic, O(1) directory renames—which is essential for Spark commit operations and avoiding partial writes.
Import mode delivers top-tier sub-second query performance by loading data into the Power BI Analysis Services cache, but requires scheduled refresh cycles and duplicates data. Direct Lake mode reads directly from Delta Parquet files in ADLS Gen2 / OneLake at in-memory speeds without data duplication and without scheduled refresh delays.
Microsoft Purview connects directly to Databricks Unity Catalog via managed metadata connectors. It automatically ingests table schemas, tags, descriptions, and column-level lineage into the enterprise Purview Data Map, ensuring centralized governance across both Azure and multi-cloud services.
A secure enterprise perimeter uses Azure Private Endpoints for all storage accounts, key vaults, and database instances. Azure Databricks clusters are deployed via VNet injection with no public IPs, and cross-service communication traverses Azure private backbone networks or ExpressRoute.
Head of Cloud & SRE Practice
Part of the Distributed Systems & Cloud Engineering at Sathus Technology. Specializing in mission-critical data lakehouses, streaming analytics, and compliance-driven platforms.
Engage Sathus certified Azure and Databricks architects to design, deploy, and govern your Azure Modern Data Platform.