Back to case studies
MICROSOFT FABRICTechnology

82% Faster Query Execution on Microsoft Fabric OneLake

Client
Global Enterprise – Technology & Analytics Division
Industry
Technology
Duration
6 months
82%
Faster query execution
97%
Less data scanned
83%
Lower compute consumption
20+
Sources unified into OneLake
The Challenge

Where things stood before

The client operated 20+ disparate data sources across Azure, AWS and on-premises SQL Server. Analytics workloads were slow (30–60 min queries), scanning terabytes for every dashboard refresh. Compute costs were spiraling and there was no unified governance layer.

Our Approach

How we engineered the outcome

  • Assessed all 20+ source systems and mapped a unification blueprint to Microsoft Fabric OneLake.
  • Implemented adaptive Liquid Clustering on high-cardinality columns to accelerate filter pushdown.
  • Added incremental statistics collection and predicate pushdown across gold-layer tables.
  • Refactored PySpark transformations for parallel processing and skew mitigation.
  • Built metadata-driven ingestion pipelines to standardize onboarding for new sources.
The Outcomes

Measurable business impact

82%

Faster query execution

97%

Less data scanned

83%

Lower compute consumption

20+

Sources unified into OneLake

Architecture Diagram

Reference Architecture

The reference architecture we deployed for this engagement, layer by layer.

Sources
  • SQL Server (on-prem)
  • AWS S3
  • Azure SQL
  • SaaS APIs
Ingest
  • Azure Data Factory
  • OneLake Shortcuts
Lakehouse
  • Microsoft Fabric OneLake
  • Delta Lake · Bronze/Silver/Gold
Transform
  • PySpark (skew-tuned)
  • dbt-fabric models
Serve
  • Fabric Warehouse
  • Power BI Direct Lake
Before / After

The transformation, in numbers

Before
Query latency30–60 min
Data scanned/query~1.2 TB
Compute per refreshF64 saturated
Sources unified0
After
Query latency<30 sec
Data scanned/query~35 GB
Compute per refresh~17% of F64
Sources unified20+