I build the plumbing, then I check the numbers that come out of it.
- 01
I started where most data people start: a table that disagreed with another table. Chasing that disagreement to its root turned out to be the whole job, and I liked it more than I expected to.
- 02
Now I write pipelines for a living — medallion layers, dedup rules, control totals. The interesting part is never the transformation. It's the check that stops a wrong number from reaching a dashboard.
- 03
The ten projects here run the same idea across three tracks: land it, model it, serve it, and let a language model take a turn at it. Each repo separates what was measured from what was not, because a portfolio number you can’t reproduce is a claim, not a result.
- 04
One README leaves a performance table deliberately blank. Filling it with plausible estimates would have been easier and would have made the project worse.
Currently
Data Engineer · ICICI Bank
- Build and operate Azure data pipelines — Data Factory and Databricks — moving customer and transaction data at bank scale.
- Work on MDM and UCIC deduplication, where two records being wrongly merged is a customer-facing failure.
- SQL across Oracle and MySQL, with PySpark for the heavier transformations.
Depth, honestly
- Data Engineeringworking depth · 82%
- Data Science / MLworking depth · 64%
- AI Engineeringloading… · 47%
Method
I lift most mornings, and progressive overload is the only learning method I trust: add a small amount of load, keep the form honest, log it, repeat. Projects work the same way — one harder constraint at a time, written down, and no credit for a rep you didn’t finish.