From batch to streaming.
Worked across legacy ETL systems and a new Azure streaming platform, from proof of concept to production.
More info
Data Engineer / Software Engineer
India · October 2020 – September 2023
Azure streaming platform
Brought multiple data sources into a shared streaming platform.
- Built PySpark pipelines to ingest and standardize JSON events, contributing to a shared schema across data sources.
- Combined streaming with batch retries for failed events and enriched orders with PostgreSQL customer data.
- Made Spark SQL transformations configurable so business logic could change without code deployments.
- PySpark
- Spark SQL
- Event Hub
- Kafka
- ADLS Gen2
- PostgreSQL
Legacy ETL & production support
Supported commission calculations on roughly 30-minute batch cycles.
- Maintained Informatica and Oracle workflows, adding transformations and loading logic for new files and tables.
- Validated incoming files with Unix scripts for checksums and date integrity before ingestion.
- Resolved production pipeline failures in coordination with the operations team.
- Informatica PowerCenter
- Oracle
- Unix
Integration & orchestration
Connected independently developed modules into one end-to-end system.
- Led pipeline integration and coordinated code changes across teams.
- Orchestrated batch and streaming workloads in Azure Data Factory, including schedules, dependencies, and automated triggers.
- Azure Data Factory
- PySpark

