Pune, India — Open to remote & hybrid

I build the pipelines that hold up
after the demo is over

Data Engineer with 4+ years turning messy, multi-source data into batch pipelines that run on schedule — across Azure, GCP and Databricks.

4+
years in production data engineering
4
platforms — Azure · GCP · Databricks · Hive
1
peer-reviewed publication, Scopus-indexed

Work

  1. May 2026 — Present Pune, India

    Data Engineer · Syngenta

    • Migrating enterprise data tables from Amazon Redshift to Databricks Unity Catalog to modernize the data platform.
    • Implementing PII tokenization for sensitive data fields in Databricks using Skyflow, strengthening data privacy and regulatory compliance posture.
  2. Dec 2023 — May 2026 Gurgaon, India

    Data Engineer · Concentrix

    • Migrating legacy pipelines off Azure and Pentaho onto GCP with Apache Spark, for better scalability, performance and cost.
    • Built CI/CD for data engineering projects using GitLab and dbt.
    • Synced Unity Catalog with Hive Metastore so tables stay visible inside the Dremio Hive catalog.
    • Built end-to-end batch pipelines in PySpark, SQL and Python to consolidate multi-source data.
    • Replaced a manual, repetitive workflow with an RPA solution built on Selenium.
  3. Sep 2022 — Nov 2023 Gurgaon, India

    Associate Data Engineer · WebHelp

    • Monitored and analyzed Azure Data Factory pipelines in production.
    • Led the migration of a reporting application's pipelines to a newer version via Azure DevOps.
  4. Aug 2021 — Sep 2022 Ropar, India

    Project Associate · IIT Ropar

    • Built a custom Yocto Linux image for streamlined software updates.
    • Implemented OTA updates at the application and system level with Mender, saving 100+ hours of manual effort.
  5. Dec 2020 — Jun 2021 Mandi, India

    Project Associate · IIT Mandi

    • Interpolated weather data from a 1° to 0.5° grid to sharpen a prediction model's accuracy.
    • Cleaned farmer crop data in Excel to improve prediction quality shown to end users.

Stack

Cloud & Orchestration

DatabricksADLSADFCloud StorageComposerDataProcAirflow

Big Data

SparkPySparkSpark SQL

Languages

PythonSQL

Engineering

Azure DevOpsGitData ModelingETL/ELT

Also using

dbtFlaskJiraConfluence

Research

Scopus-indexed

Optimized ensemble machine learning framework for high dimensional imbalanced bioassays

Sharma, R., Hooda, N. (2019) — Revue d'Intelligence Artificielle, Vol. 33, No. 5, pp. 387–392.

Read the paper ↗

Education

Aug 2018 — Jun 2020

M.E., Computer Science & Engineering

Chandigarh University, Mohali, India

Thesis: Optimized ensemble machine learning framework for high dimensional imbalanced bioassays.

Aug 2013 — May 2017

B.Tech, Computer Science & Engineering

Career Point University, Hamirpur, India

Contact

// let's build something

Migrating a legacy pipeline, standing up a lakehouse, or just want to talk Spark and Databricks?
My inbox is open.

in LinkedIn </> GitHub
Pune, India · IST (UTC+5:30) ~ 24h response time