Pune, India — Open to remote & hybrid
I build the pipelines that hold up
after the demo is over
Data Engineer with 4+ years turning messy, multi-source data into batch pipelines that run on schedule — across Azure, GCP and Databricks.
- 4+
- years in production data engineering
- 4
- platforms — Azure · GCP · Databricks · Hive
- 1
- peer-reviewed publication, Scopus-indexed
// this is roughly how the pipeline actually runs
Work
-
Data Engineer · Syngenta
- Migrating enterprise data tables from Amazon Redshift to Databricks Unity Catalog to modernize the data platform.
- Implementing PII tokenization for sensitive data fields in Databricks using Skyflow, strengthening data privacy and regulatory compliance posture.
-
Data Engineer · Concentrix
- Migrating legacy pipelines off Azure and Pentaho onto GCP with Apache Spark, for better scalability, performance and cost.
- Built CI/CD for data engineering projects using GitLab and dbt.
- Synced Unity Catalog with Hive Metastore so tables stay visible inside the Dremio Hive catalog.
- Built end-to-end batch pipelines in PySpark, SQL and Python to consolidate multi-source data.
- Replaced a manual, repetitive workflow with an RPA solution built on Selenium.
-
Associate Data Engineer · WebHelp
- Monitored and analyzed Azure Data Factory pipelines in production.
- Led the migration of a reporting application's pipelines to a newer version via Azure DevOps.
-
Project Associate · IIT Ropar
- Built a custom Yocto Linux image for streamlined software updates.
- Implemented OTA updates at the application and system level with Mender, saving 100+ hours of manual effort.
-
Project Associate · IIT Mandi
- Interpolated weather data from a 1° to 0.5° grid to sharpen a prediction model's accuracy.
- Cleaned farmer crop data in Excel to improve prediction quality shown to end users.
Stack
Cloud & Orchestration
DatabricksADLSADFCloud StorageComposerDataProcAirflow
Big Data
SparkPySparkSpark SQL
Languages
PythonSQL
Engineering
Azure DevOpsGitData ModelingETL/ELT
Also using
dbtFlaskJiraConfluence
Research
Scopus-indexed
Optimized ensemble machine learning framework for high dimensional imbalanced bioassays
Read the paper ↗Education
Aug 2018 — Jun 2020
M.E., Computer Science & Engineering
Chandigarh University, Mohali, India
Thesis: Optimized ensemble machine learning framework for high dimensional imbalanced bioassays.
Aug 2013 — May 2017
B.Tech, Computer Science & Engineering
Career Point University, Hamirpur, India