// data engineer · aws · databricks · pyspark · airflow
ETL/ELT systems, serverless AWS workflows, Databricks medallion pipelines, and automation that lets teams trust their reporting and stop doing manual work.
The short version, written for decision-makers. If you're an engineer and want the diagrams, service choices, and design rationale — the full technical case studies are one click away.
An organization can't report on data it can't trust. This platform turns raw transactional extracts into quality-checked, audit-ready reporting layers — refreshed incrementally, not rebuilt from scratch.
A shared analytics platform serving research and commercial teams across the business. My focus: keeping a large estate of automated workflows healthy, finding root causes fast, and delivering changes to production safely.
Every new data provider means a new feed with its own quirks. This system made adding the next feed cheap, fast, and safe — fully serverless, so infrastructure cost scales with actual usage, not idle servers.
Where it started: Python automation engineering at AlphaSol — repetitive client work turned into systems that saved 30+ hours of manual effort per month.
Details →Who needs the data, what decision it supports, freshness needs, failure impact.
APIs, files, schemas, credentials, volume, latency, and ownership — before coding.
Cloud-native ingestion, transformation, storage, orchestration, retry patterns.
Quality checks positioned at layer boundaries — not scripts bolted on after the fact.
Runbooks with design, implementation, and rollback — plus mentoring teammates on the patterns.
PMAS-Arid Agriculture University Rawalpindi (2018–2022). Final-year project: a number-plate recognition system built with YOLOv4, Flask, MariaDB, frontend web technologies, and Google Vision API experiments.
Data engineering roles, remote contracts, and cloud ETL projects — let's design and deliver the system together.