Skip to content
View hamzaali026's full-sized avatar
πŸ’­
🟒 Open to Work | Data Engineering | Azure | Databricks | Snowflake
πŸ’­
🟒 Open to Work | Data Engineering | Azure | Databricks | Snowflake

Block or report hamzaali026

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
hamzaali026/README.md

Hi, I'm Hamza Ali πŸ‘‹

Senior Data Engineer | Cloud Data Platforms | Healthcare Data | Azure | Microsoft Fabric

I'm a Senior Data Engineer with 8+ years of experience designing, building, and optimizing scalable data platforms and production-grade data pipelines across healthcare, fintech, and retail.

My work focuses on turning complex data requirements into reliable, scalable, secure, and cost-efficient data products. I specialize in the Microsoft Azure and Microsoft Fabric ecosystems, with extensive hands-on experience across Databricks, Snowflake, dbt, Apache Spark, Kafka, and cloud-native data architectures.

I enjoy solving challenging data problemsβ€”from building high-volume batch and real-time pipelines to designing governed lakehouses and modernizing legacy data platforms.


πŸš€ What I Do

  • πŸ—οΈ Design and architect modern cloud data platforms
  • ⚑ Build high-performance batch and real-time data pipelines
  • ☁️ Develop cloud solutions across Azure, AWS, and GCP
  • πŸ₯ Build and integrate healthcare data platforms using HL7, FHIR, and EHR/EMR data
  • 🧱 Design Lakehouse, Medallion, Data Vault, and dimensional architectures
  • πŸ”„ Develop scalable ETL/ELT pipelines with dbt, Spark, ADF, and Airflow
  • πŸ“Š Build analytics solutions with Power BI and Tableau
  • πŸ” Implement data governance, security, masking, and compliance
  • πŸš€ Automate infrastructure and deployments using Terraform and CI/CD
  • πŸ“ˆ Optimize data platforms for performance, reliability, and cost

πŸ› οΈ Tech Stack

☁️ Cloud & Data Platforms

Microsoft Azure Microsoft Fabric AWS GCP Azure Data Factory Azure Synapse OneLake Databricks S3 Glue Lambda BigQuery Dataflow

🏒 Data Warehousing & Lakehouse

Snowflake Databricks Delta Lake Microsoft Fabric Apache Iceberg Medallion Architecture Data Vault 2.0 Star Schema Snowflake Schema

πŸ”„ Data Engineering

Apache Spark PySpark dbt Apache Airflow Apache Kafka Spark Structured Streaming Apache Beam ETL ELT Batch Processing Real-Time Processing

πŸ’» Programming

Python SQL T-SQL PL/pgSQL Scala JavaScript TypeScript

πŸ—„οΈ Databases & Storage

PostgreSQL MySQL SQL Server Aurora Cassandra DynamoDB Redis HBase ClickHouse

πŸ₯ Healthcare Data

HL7 v2 FHIR EHR/EMR Integration HIPAA PHI De-identification Data Masking Immuta Data Lineage Data Governance

πŸ“Š Analytics & BI

Power BI Tableau Semantic Models KPI Dashboards Executive Reporting

βš™οΈ DevOps & Reliability

Terraform GitHub Actions Azure DevOps Great Expectations DataOps Datadog New Relic CI/CD Monitoring Performance Tuning


πŸ’Ό Professional Experience

πŸ₯ Senior Data Engineer β€” People Inc.

Jan 2022 – Present

  • Architected a HIPAA-compliant Azure Databricks + Delta Lake lakehouse using Medallion Architecture.
  • Standardized ingestion of HL7 v2 and FHIR clinical feeds from EHR systems.
  • Improved downstream query performance by 40% while reducing storage costs by 30%.
  • Led migration of analytics workloads to Microsoft Fabric and OneLake, supporting 200+ business and clinical users.
  • Built production-grade dbt + Snowflake ELT pipelines with automated testing and CI/CD.
  • Designed real-time pipelines using Kafka, Spark Structured Streaming, and Databricks.
  • Implemented data governance and privacy controls using Immuta, including PHI de-identification and dynamic masking.
  • Developed AWS ingestion pipelines processing 5TB+ of raw data daily.
  • Automated multi-cloud infrastructure using Terraform.
  • Increased pipeline reliability to 99.9% through data-quality testing, monitoring, alerting, and automated retries.
  • Partnered with product, compliance, analytics, and executive stakeholders to deliver data-driven solutions.

πŸ›’ Data Engineer β€” DataSap Inc.

Mar 2019 – Dec 2021

  • Built cloud-native data pipelines using Azure Data Factory, Databricks, and Azure Synapse.
  • Migrated legacy on-premises ETL workloads to cloud infrastructure.
  • Developed real-time pipelines using Kafka and PySpark for e-commerce analytics.
  • Migrated retail reporting workloads to Snowflake, reducing reporting time from 4 hours to under 15 minutes.
  • Designed Star, Snowflake, and Data Vault data models.
  • Automated infrastructure deployment with Terraform, reducing cloud provisioning errors by 80%.
  • Developed Power BI dashboards for operational and business KPI reporting.
  • Worked within Agile/Scrum teams to deliver production data solutions.

πŸ’» Junior Data Engineer β€” Infosys

Jun 2017 – Feb 2019

  • Developed ETL pipelines using Python and SQL.
  • Integrated data from REST APIs, flat files, and relational databases.
  • Designed PostgreSQL and MySQL schemas and optimized SQL queries and stored procedures.
  • Contributed to AWS migrations using RDS, EC2, S3, and IAM.
  • Built automated data cleansing and transformation workflows.
  • Created data dictionaries and technical documentation.

πŸ“ˆ Selected Impact

Achievement Impact
Pipeline processing optimization 65% faster
Query performance improvement 40% faster
Storage cost reduction 30%
Daily AWS data processing 5TB+
Snowflake reporting optimization 4 hours β†’ <15 minutes
Cloud provisioning error reduction 80%
Production pipeline reliability 99.9%
Real-time lakehouse processing 50M+ events/day
Infrastructure cost savings ~$120K/year
Governed data platform users 200+

🧠 Architecture & Engineering Focus

I particularly enjoy working on architectures such as:

                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚   Data Sources       β”‚
                    β”‚                      β”‚
                    β”‚ EHR / HL7 / FHIR     β”‚
                    β”‚ APIs / Applications  β”‚
                    β”‚ Databases / Events   β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚   Data Ingestion     β”‚
                    β”‚                      β”‚
                    β”‚ ADF / Kafka / Glue   β”‚
                    β”‚ APIs / Streaming     β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β”‚       Cloud Data Platform       β”‚
              β”‚                                 β”‚
              β”‚ Azure / Fabric / Databricks     β”‚
              β”‚ Delta Lake / OneLake / Snowflakeβ”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                              β”‚
                              β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚ Transformation       β”‚
                    β”‚                      β”‚
                    β”‚ Spark / PySpark/dbt β”‚
                    β”‚ ETL / ELT            β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚ Analytics & AI       β”‚
                    β”‚                      β”‚
                    β”‚ Power BI / Tableau   β”‚
                    β”‚ ML / Data Products   β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ” Data Engineering Principles

I care deeply about building data platforms that are:

Reliable β€” automated testing, monitoring, retries, and strong SLAs

Scalable β€” designed to handle growing data volumes and workloads

Secure β€” governed access, masking, de-identification, and compliance

Maintainable β€” modular code, documentation, version control, and CI/CD

Cost-efficient β€” continuous performance and infrastructure optimization

Business-focused β€” engineering solutions that produce measurable outcomes


🌎 Areas of Expertise

  • Cloud Data Engineering
  • Healthcare Data Engineering
  • Data Platform Architecture
  • Lakehouse Architecture
  • Real-Time Data Streaming
  • Data Warehousing
  • ETL / ELT
  • Data Governance & Security
  • Data Quality & Observability
  • Infrastructure as Code
  • Analytics Engineering
  • Business Intelligence

πŸ“š Currently Focused On

Building and improving modern cloud-native data platforms, with a particular interest in:

  • Microsoft Fabric & OneLake
  • Lakehouse Architecture
  • Databricks & Delta Lake
  • Snowflake
  • Real-Time Streaming
  • Data Governance
  • Healthcare Data Platforms
  • Scalable Data Products

🀝 Let's Connect

I'm always interested in connecting with engineers, architects, data professionals, and teams working on interesting problems in data engineering, cloud platforms, analytics, and healthcare technology.

πŸ“§ Email: hamzaa.ali026@gmail.com πŸ’Ό LinkedIn: [linkedin.com/in/hamzaa-ali-data/


⭐ If you find my projects useful, consider giving them a star!

Thanks for visiting my profile!

Popular repositories Loading

  1. hamzaali026 hamzaali026 Public

  2. ClinLake-Nexus ClinLake-Nexus Public

    A production-inspired healthcare lakehouse that transforms FHIR clinical data into governed, analytics-ready datasets using PySpark, Delta Lake, Medallion Architecture, and Databricks.

    Python

  3. FHIRBridge-Atlas FHIRBridge-Atlas Public

    A healthcare interoperability pipeline that validates, normalizes, and transforms FHIR resources into analytics-ready clinical datasets with schema validation, quality controls, and scalable proces…

    Python

  4. Clinora-Intelligence Clinora-Intelligence Public

    A production-inspired clinical intelligence platform that transforms healthcare data into trusted analytics and AI-ready datasets using PySpark, Delta Lake, FHIR, Medallion Architecture, and ML.

    Python

  5. ClinSight-Analytics ClinSight-Analytics Public

    A modern clinical analytics platform that transforms FHIR and healthcare data into governed patient insights, quality metrics, and reusable analytical datasets using Python, Spark, and SQL.

    Java

  6. CareStream-Fabric CareStream-Fabric Public

    A scalable healthcare data pipeline that ingests clinical and operational data, applies quality and transformation workflows, and delivers analytics-ready datasets through a cloud-native lakehouse …

    Python