From batch to streaming.
Worked across legacy ETL systems and a new Azure streaming platform, from proof of concept to production. Helped reduce processing latency from 20–30 minutes to about 10 seconds.
More info
Data Engineer / Software Engineer
India · October 2020 – September 2023
Modern Azure data platform
- Designed and developed PySpark streaming pipelines to ingest, transform, and standardize JSON data from multiple sources using Azure Event Hub, Kafka, and ADLS Gen2.
- Helped design a hybrid architecture combining near-real-time streaming with batch retries for failed events, reducing end-to-end latency from 20–30 minutes to about 10 seconds.
- Contributed to a team-designed standardized JSON schema, allowing different data sources to flow through a shared pipeline.
- Implemented configuration-driven Spark SQL transformations so queries and business logic could change without code deployments.
- Enriched streaming order events with PostgreSQL customer master data, delivering complete records downstream.
- Led the integration of independently developed pipeline modules into one end-to-end system, coordinating code changes across teams.
- Orchestrated batch and streaming workloads with Azure Data Factory, including scheduling, dependencies, and automated triggers.
Legacy ETL & production support
- Maintained Informatica PowerCenter ETL workflows and Oracle batch systems for commission calculations, running on roughly 30-minute cycles.
- Developed Unix shell scripts to validate incoming files, verify checksums, and check date integrity before ingestion.
- Extended existing workflows to onboard new sources, adding field-level transformations and loading logic for new files and tables.
- Troubleshot production pipeline failures and coordinated their resolution with the operations team.
PySpark · Spark SQL · Azure Event Hub · Kafka · ADLS Gen2 · Azure Data Factory · PostgreSQL · Informatica · Oracle · Unix

