Skip to main content

Command Palette

Search for a command to run...

Global Data Residency Platform - Architecture Explained

Updated
•12 min read•View as Markdown
Global Data Residency Platform - Architecture Explained
S
Sharing enterprise grade data engineering case studies, cloud architectures, and production lessons from the real world.

Category: Data Engineering · Data Platform · Data Governance · Cloud Architecture

1. Business Problem

A multinational enterprise operates across multiple regions:

  • EU Europe

  • US United States

  • APAC Asia Pacific

Each region stores customer information, for example:

  • Customer Name

  • Email

  • Phone Number

  • Address

  • Purchase History

  • Payment Details

Each region is bound by a different privacy regulation:

Region Regulation Requirement
Europe GDPR Customer data cannot leave the EU
United States CCPA Customer data must remain compliant with US privacy laws
APAC PDPA Customer data may be restricted from leaving specific countries

Example: A customer in Germany purchases a product through an e-commerce platform. According to GDPR, that customer's data cannot be replicated into a US-based data lake - even for internal reporting.

The challenge:

Customer data should never leave its legal region, but executives still need a single, global view of the business - revenue, growth, inventory, forecast accuracy - across all regions.


2. The Naive Approach (and Why It Fails)

Most engineers' first instinct is:

Create one ADLS storage account
Create three containers: eu / us / apac
Apply RBAC per container
Done.

This looks correct, but it confuses two very different concerns.

A Storage Account Has One Physical Location

When an Azure Storage Account is created, Azure requires a single region e.g. West Europe, East US, or Australia East. Every container inside that account is physically stored in that one location.

Storage Account (Location = East US)
├── eu
├── us
└── apac

Even with perfect RBAC, European customer data is still physically sitting in the United States. For GDPR, this is typically unacceptable.

RBAC ≠ Data Residency

These are two completely different questions:

Concept Question it answers Example
RBAC / Access Control Who can access the data? "Only the EU team can read the eu container"
Data Residency Where is the data physically stored? "Storage Account location = East US"

Restricting who can read data does nothing to change where that data physically lives. A single shared storage account - no matter how well access is locked down - cannot satisfy a hard residency requirement like GDPR's.


Instead of using a single shared storage account, the platform provisions a fully independent data platform for each region. Each regional platform has its own subscription, storage, compute, governance, and security services.

Component Europe (EU) United States (US) APAC
Azure Subscription Dedicated Dedicated Dedicated
Azure Data Factory Regional Instance Regional Instance Regional Instance
ADLS Gen2 Regional Storage Regional Storage Regional Storage
Medallion Architecture Bronze / Silver / Gold Bronze / Silver / Gold Bronze / Silver / Gold
Azure Databricks Regional Workspace Regional Workspace Regional Workspace
Unity Catalog Regional Governance Regional Governance Regional Governance
Microsoft Purview Enterprise Metadata & Lineage Enterprise Metadata & Lineage Enterprise Metadata & Lineage

Each region operates independently, ensuring that customer data remains within its legal boundary while maintaining a consistent platform architecture across the enterprise.

Every region owns its own:

  • Storage - physically located inside that legal jurisdiction

  • Compute - Databricks clusters run in-region

  • Governance - Unity Catalog + Purview scoped per region

  • Metadata - no cross-region metadata dependency for raw data

  • Security - independent IAM, networking, and audit trail

No raw customer data is ever shared between regions. Only aggregated, PII-free business metrics are shared globally.

Global Multi-Region Architecture Overview


4. Architecture Components

4.1 Regional Data Platform

Every region has a fully independent Azure environment following the same pattern:

Source Systems → Azure Data Factory → ADLS Gen2 → Bronze → Silver → Gold

The same architecture is deployed independently in the EU, US, and APAC - there is no shared control plane.

4.2 Data Ingestion

Azure Data Factory ingests data from regional source systems, for example:

  • SAP

  • Salesforce

  • CRM / ERP / POS systems

  • Amazon / e-commerce applications

Data never crosses regional boundaries during ingestion - each ADF instance only talks to source systems inside its own region.

4.3 Storage Layer - Medallion Architecture

Azure Data Lake Storage Gen2 stores data using the Bronze → Silver → Gold pattern:

Layer Purpose
Bronze Raw, unprocessed data as landed from source systems
Silver Validated, cleaned, deduplicated data
Gold Business-ready, curated data ready for consumption

Each region maintains its own storage account, physically located inside that region.

4.4 Data Processing

Azure Databricks processes data locally within each region. Typical transformations:

  • Cleaning & standardization

  • Deduplication

  • Business rule application

  • Data validation

  • Delta Lake optimization (ACID transactions, time travel)

Output is written into the regional Gold tables.

4.5 Data Aggregation

Instead of sending customer-level records globally:

❌ Not Allowed

Customer Email Revenue
John john@gmail.com $500

✅ Allowed - Aggregated Metrics Only

Country Revenue Orders Top Product
Germany $12M 1.4M Pedigree
France $8M 0.9M -

No personal information ever leaves the region.

4.6 Global Reporting Layer

The global layer receives only business metrics:

  • Revenue, Sales, Orders

  • Monthly Growth

  • Inventory & Forecast

  • Product Performance

This enables global dashboards and executive reporting without violating any regional data residency regulation, because Power BI never connects to raw customer data - it only consumes the aggregated dataset each regional platform produces.

Detailed Regional Pipeline (Source → Storage → Processing → Governance → Aggregation → Global Layer)


5. Why Two Different Governance Tools? Purview vs. Unity Catalog

A common question: "We already have Microsoft Purview, why do we also need Unity Catalog?"

They solve different problems and operate at different layers.

Microsoft Purview - Enterprise Metadata & Discovery

Purview is an enterprise-wide metadata platform. It scans across many systems:

SAP · Azure SQL · Oracle · Snowflake · Databricks · Power BI · Blob Storage

It answers questions like:

  • Where is Customer Email stored across the company?

  • Which datasets contain PII?

  • Which reports consume this table?

  • What is the end-to-end data lineage?

Purview provides visibility. It does not enforce runtime permissions inside Databricks.

Unity Catalog - Databricks-Native Runtime Governance

Unity Catalog governs access inside Databricks. It answers:

SELECT * FROM gold.customer

Can this user run this query? Should a row be filtered? Should a column be masked?

Unity Catalog performs real-time authorization decisions:

  • Allow / Deny

  • Row-level filters

  • Column-level masking (e.g., hide salary, mask email)

Side-by-Side Comparison

Capability Unity Catalog Microsoft Purview
Databricks access control ✅ ❌
Catalogs & schemas ✅ ❌
Row-level security ✅ ❌
Column masking ✅ ❌
Table permissions ✅ ❌
Databricks-native lineage ✅ Partial
Enterprise-wide metadata ❌ ✅
Scans SAP / Azure SQL / Snowflake ❌ ✅
Scans Power BI ❌ ✅
Enterprise data discovery ❌ ✅

How They Work Together - Example Pipeline

SAP → Azure Data Factory → ADLS → Databricks → Power BI
Step Purview Unity Catalog
1. Purview scans SAP Discovers Customer Email, Customer ID, flags PII -
2. ADF copies data Records lineage -
3. Databricks processes data - Controls who can read / write / modify / create tables
4. Power BI reads aggregated metrics Extends lineage to dashboards Already enforced access before Power BI consumed the data

Two Layers, Two Jobs

Enterprise Layer  →  Microsoft Purview  →  Knows "what data exists?"
Platform Layer    →  Unity Catalog      →  Controls "who can access it?"
Analytics Layer   →  Power BI           →  Consumes only aggregated metrics — never raw customer data

6. Complete Data Flow

Business Applications (SAP, CRM, ERP, POS, Salesforce, Amazon)
                    │
                    ▼
          Azure Data Factory
                    │
   ═══════════════════════════════════════════
   EU Platform          US Platform        APAC Platform
   ADLS → Bronze →      (same flow)        (same flow)
   Silver → Gold
   → Purview scan
   → Unity Catalog auth
   → Regional Aggregation
   ═══════════════════════════════════════════
                    │
                    ▼
          Global Analytics Dataset
        (Aggregated Metrics Only — No PII)
                    │
                    ▼
             Power BI Dashboards
               Executive Users

7. Technology Responsibilities

Technology Responsibility
Azure Data Factory Region-scoped data ingestion
ADLS Gen2 Regional storage (physically resident in-region)
Azure Databricks In-region data processing
Delta Lake ACID storage & time travel
Unity Catalog Databricks runtime authorization, row/column security
Microsoft Purview Enterprise-wide metadata, classification, lineage
Power BI Executive dashboards on aggregated metrics only

8. Example Scenario

Europe

Name: John | Email: john@gmail.com | Country: Germany | Revenue: $500

Stored only inside Europe.

United States

Customer: Mike | Country: Texas | Revenue: $700

Stored only inside the US.

APAC

Customer: Ravi | Country: India | Revenue: $400

Stored only inside APAC.

Global Layer Receives

Europe Revenue: $2.1B     US Revenue: $3.5B     APAC Revenue: $1.8B

No customer information exists at this layer.


9. Benefits of the Regional Platform Approach

Benefit Why it matters
Regulatory Compliance GDPR, CCPA, and PDPA compliant by design - customer data never leaves its legal jurisdiction
True Data Residency Storage is physically located in-region, not just access-restricted
Performance Regional workloads stay local - no unnecessary cross-region traffic
Disaster Isolation An EU outage does not affect US or APAC operations
Security Each region has fully independent IAM, storage, compute, and networking
Independent Scaling Europe can scale without impacting APAC or US workloads
Centralized Analytics Executives still get one unified global view, built entirely from aggregated, PII-free metrics

10. Why This Is an Expert - Level Case Study

This problem combines three complex engineering domains:

  1. Distributed Data Platforms - multiple regional platforms operating fully independently, each with its own storage, compute, and security boundary.

  2. Enterprise Data Governance - centralized enterprise metadata (Purview) layered on top of decentralized, runtime-enforced access control (Unity Catalog).

  3. Global Business Analytics - executive reporting built entirely from aggregated metrics, with zero exposure of sensitive customer information.

Designing such a platform requires expertise in:

  • Azure Databricks & Delta Lake

  • Data governance and data residency law

  • Enterprise security architecture

  • Cloud architecture & distributed data engineering

This is a common pattern for global enterprises in regulated industries: banking, healthcare, retail, pharmaceuticals, and technology.


11. Key Takeaways

Approach Verdict
One ADLS account with regional containers + RBAC Technically possible, but fails true data residency - data has one physical location regardless of access rules
Fully separate regional platforms (this design) Satisfies compliance, performance, security, disaster isolation, and independent scaling
Microsoft Purview Enterprise-wide metadata and governance - knows what data exists across all systems
Unity Catalog Databricks-native governance controls who can access data inside Databricks, at runtime
Power BI Never touches raw customer data - consumes only the aggregated business metrics each region produces

Appendix: Architecture Diagrams

This case study includes two architecture diagrams to help visualize the solution.

Diagram Description
Diagram 1 High-level architecture showing separate regional data platforms (EU, US, and APAC) with customer data remaining within each region, while only aggregated, non-PII business metrics flow into a centralized global reporting layer for Power BI.
Diagram 2 Detailed view of a single regional platform illustrating the complete data flow from source systems through ingestion, storage (Bronze/Silver/Gold), Databricks processing, Unity Catalog governance, regional aggregation, and finally publishing aggregated metrics to the global analytics layer.

More from this blog

D

Data Engineering

8 posts

Real-world Data Engineering case studies, enterprise architectures, cloud platforms, and production-ready solutions for modern data teams.