Open to Data & ML Engineering roles · Bangalore, India

I build data systems
that hold up to scrutiny.

I'm a data engineer, two years into building production pipelines on Azure Databricks and AWS. PySpark, Delta Lake, star-schema Gold models, and the Power BI semantic layers that sit on top of them. Outside work I build and deploy machine learning and LLM-backed services in Python. Most of that time goes on evaluation.

Experience

Production data engineering for enterprise clients.

Two years at Accenture on Azure Databricks, preceded by two internships in cloud and database engineering.

Sep 2024 — Present

Data Engineer — Packaged App Development Associate

Accenture · Bangalore

  • Built Bronze → Silver → Gold Delta pipelines on Azure Databricks using config-driven YAML with PySpark and Spark SQL, replacing per-source hardcoded jobs with one reusable ingestion pattern across an industrial IoT programme and a UK financial-services client.
  • Designed star-schema Gold models and unified heterogeneous per-equipment tag schemas through canonical renaming and unionByName, giving downstream consumers one conformed grain instead of raw per-equipment tables.
  • Diagnosed and resolved a source-table migration serving stale data, then built validation notebooks enforcing grain, reconciliation and null-threshold checks to block bad loads from reaching Gold.
  • Delivered enterprise Power BI semantic models with reusable DAX measure libraries and dynamic row-level security through Entra ID, from requirement through UAT into production.
  • Translated client SME reporting requirements into pipeline and model design decisions, and documented pipeline logic and data definitions to support handover.
Feb 2023 — Aug 2023

Database Management Intern

T-Web Exponent Services

  • Optimised relational and NoSQL databases and tuned SQL queries to improve query performance, and integrated backend APIs for low-latency data flow between services.
  • Migrated on-premise database infrastructure to cloud, improving reliability and uptime.
Jul 2022 — Jan 2023

AWS Solutions Intern

T-Web Exponent Services

  • Engineered AWS Glue ETL jobs automating data transforms, improving throughput 20% over a three-month project.
  • Deployed S3, EC2, Lambda and Athena, and implemented IAM roles giving services controlled cross-service access.

Selected work

Four systems, built end to end.

Ingestion, modelling, evaluation, deployment. All four are running, and the links below open the real thing rather than a screenshot.

01

ReportScope

LivePharmacovigilance

Drug-safety signal detection over the FDA adverse-event database

A pipeline and lookup tool over the openFDA adverse-event database. It flags drug–reaction pairs worth a human looking at. Every score is a closed-form statistic over a 2×2 table, and each one carries its own bias diagnostics.

33,852scored drug–reaction pairs
361drugs · 305 molecules
20.7MFAERS reports scanned
13 / 13known-answer controls pass
  • BCPNN and MGPS disproportionality scoring with empirical-Bayes shrinkage. Five hyperparameters fitted by MLE across all 33,852 cells, so the prior comes from the data.
  • A known-answer harness of 9 positive and 4 negative controls. It reproduces the statin–rhabdomyolysis reporting odds ratio of 12.95, and it picked out cerivastatin, withdrawn in 2001, as the highest-scoring statin without being told to.
  • An active-comparator diagnostic that separates molecule effects from delivery artifacts. Exenatide device leakage drops from 36.23 to 1.01 against the right control. The signal was the injector, not the drug.
  • Runs as a containerised Flask service with PostgreSQL, Alembic migrations, accounts, email verification, and a weekly digest on GitHub Actions.
PythonpandasNumPySciPyFlaskPostgreSQLDockerAlembicopenFDA APIGitHub Actions
02

Cross-Patient Seizure Detection

LiveClinical EEGMachine Learning

Machine learning on the CHB-MIT scalp EEG corpus

Detecting epileptic seizures in scalp EEG for patients the model has never seen. The classifier was never the hard part. Building an evaluation you can trust is, when positives are three in a thousand and no two brains look alike. The live app runs the whole pipeline in the browser on a bundled recording, so there is nothing to upload.

980 hof recordings processed
23patients, held out one at a time
352,742labelled windows
142tests · 44 binding rules
  • Leave-one-subject-out cross-validation across 23 patients, reaching 0.31 PR-AUC against a 0.004 baseline. Getting the protocol wrong is worth anywhere from −0.10 to +0.40 depending on the model, so there is no one number to quote for a dataset.
  • Two metrics that quietly flatter a model. Averaging precision across folds read 0.23 where the pooled figure was 0.004, so metrics are pooled from summed confusion counts now. And a DummyClassifier that alarms almost constantly still scores a respectable 0.89 alarms/hour, because that metric counts alarm runs rather than time spent alarming.
  • Decision threshold and run-length cannot be tuned separately. Chosen one after the other, the pair detected nothing at all. They are searched jointly now.
  • Three upstream data bugs that corrupt labels without throwing an error: a library silently renaming duplicate channels, one mislabelled seizure end time, and an annotation field that turns out to be a byte offset rather than a timestamp.
Pythonscikit-learnHistGradientBoostingRandom ForestNumPySciPyMNEWFDBParquetpytest
03

ScoutIQ

LiveSports Analytics

Football match outcome and scoreline prediction

A match prediction engine for top-tier football. Elo supplies the outcome probabilities and a Dixon-Coles grid is rescaled onto them, so the scorelines, expected goals and over/under numbers can never contradict the match result.

3,504held-out test matches
0.9932Elo outcome log loss
0.025measured gap to the market
Dailyautomated data refresh
  • Benchmarked Elo, Dixon-Coles and XGBoost on a held-out season, picking on log loss with paired bootstrap tests instead of point estimates. A 46-feature XGBoost turned out to be statistically indistinguishable from a 2-parameter Elo.
  • Blending bookmaker odds adds nothing. Linear and logarithmic pooling both selected weight 0.00 on validation. The odds are still in the UI, they just do not touch the prediction.
  • Serves from a FastAPI backend with a Vercel frontend, refreshed by a GitHub Actions job that rebuilds Gold data and commits it back to trigger a redeploy.
PythonXGBoostscikit-learnMLflowFastAPIDatabricksDelta LakePower BIGitHub ActionsVercel
04

Compass

LiveLLM Application

AI job-search copilot with a provider-agnostic model layer

A copilot that scores how well a role fits a real CV, shows the gaps, and drafts the application. It does not scrape job boards and it does not submit anything on your behalf. The job market has a volume problem already. The live link opens a sign-in page because this is a real multi-user app with per-user isolation and cost caps, not a shared demo.

3interchangeable LLM providers
5constraints enforced by tests
1outbound call site, asserted
  • One LLM interface over Anthropic, Gemini and Ollama, so the app moves between hosted and fully local models without touching a call site. Useful when a free tier quietly drops a capability on you.
  • Job-fit scoring runs on local sentence embeddings, with a lexical fallback for when the model is not there. The hosted demo runs the fallback.
  • The product constraints are five tests. The suite fails on a banned scraping library, a restricted-platform URL, a submit-shaped route, or a model call inside the quality gate.
  • One test counts the outbound call sites in the discovery package and asserts there is exactly one. Another pins the allowlist that site checks against.
  • Parses and rewrites résumés across PDF and DOCX with layout introspection, exporting ATS-safe documents.
PythonFastAPIHTMXSQLAlchemysentence-transformersAnthropicGeminiOllamaSQLitepytest

Toolkit

What I reach for.

Bold groups are what I use daily and would be happy to be tested on cold. The rest I have worked with and can discuss honestly, including where my experience stops.

Languages

PythonSQLJavaJavaScript

Data Engineering

PySparkSpark SQLDelta LakeMedallion ArchitectureETL / ELTAirflowParquetData Quality

Machine Learning

scikit-learnXGBoostMLflowpandasNumPySciPyFeature EngineeringCross-ValidationModel Evaluation

AI & LLM

LLM API IntegrationAnthropicGeminiOllamaSentence TransformersVector EmbeddingsSemantic Search

Modelling & Warehousing

Dimensional ModellingStar SchemaGold-Layer DesignAmazon RedshiftAmazon Athena

Cloud

Azure DatabricksADLS Gen2Unity CatalogEntra IDAWS S3GlueLambdaEC2IAMRenderVercel

Backend & Databases

FastAPIFlaskREST APIsSQLAlchemyAlembicPostgreSQLMySQLNoSQL

BI & Reporting

Power BIDAXPower QueryRow-Level SecurityRBACTableau

Tools & Practices

GitGitHub ActionsCI/CDDockerpytestJIRAAgile

Education

Electronics & Communication Engineering.

2020 — 2024

B.Tech, Electronics and Communication Engineering

Narula Institute of Technology · Kolkata

Contact

Let's talk.

Open to Data Engineering and Machine Learning roles. Email is the fastest way to reach me, and I read everything.