← back to home

All projects

Everything I've built

★ Featured
Data Engineering

paytrail: Payments Medallion Lakehouse

  • Built a Databricks medallion lakehouse taking 6.3M synthetic payments from Azure storage through bronze, silver and gold layers.
  • Designed idempotent incremental loads for late and out-of-order data, with Unity Catalog governance and 51 dbt tests gated in CI.
  • Reconciled settlement outputs to source at penny level and reduced the heaviest rollup runtime by 39% after optimisation.
DatabricksDelta LakeUnity CatalogdbtAzureGitHub Actions
★ Featured
Data Engineering

Cross-System Customer Data Reconciliation

  • Reconciled customer records across Salesforce, NetSuite and Stripe-style data, starting from 12,566 supplied case rows.
  • Built deterministic exact-key cleaning for duplicate records, missing identifiers, mixed date formats and European/US currency formats.
  • Kept unresolved links visible instead of inferring them: 49 NetSuite rows missing a Stripe ID and 675 unmatched Stripe transactions surfaced as exceptions.
PythonpandasDeterministic Matching
★ Featured
Data Engineering

Workplace Safety Analytics Pipeline

  • Built a batch pipeline ingesting 688K+ OSHA injury records into a PostgreSQL star schema, orchestrated through a seven-task Airflow DAG.
  • Added local LLM enrichment for free-text incident narratives using Ollama, keeping the pipeline reproducible and inference cost-free.
  • Enforced 81 automated data and contract tests through dbt and GitHub Actions, with failed checks blocking the build.
Apache AirflowdbtPostgreSQLDockerGitHub ActionsOllamaPython
★ Featured
ML & NLP

The Edit: H&M Recommender

  • Built a two-stage recommender over 31M H&M transactions, using BigQuery SQL retrieval followed by a CatBoost ranking model.
  • Improved MAP@12 from 0.0053 to 0.0292, with a 95% bootstrap confidence interval confirming the lift was stable.
  • Added a temporal-leakage gate, diversity checks, 80+ automated tests, and a typed FastAPI prediction endpoint.
BigQueryCatBoostFastAPIPythonSQL
★ Featured
Forecasting & Research

SHAP Stability Across the Rashomon Set

  • Tested whether SHAP feature rankings remain stable when a forecasting model is replaced by another near-optimal model.
  • Ran AutoGluon and H2O experiments across six forecasting benchmarks, multiple temporal splits and Rashomon thresholds on Leiden's ALICE HPC.
  • Found H2O built larger, more varied Rashomon sets than AutoGluon (up to 15 models across 5 families, against up to 5 in 1), yet both frameworks' stability agreed within 0.04 Spearman rho on every dataset (0.77 to 0.98 overall).
AutoGluonH2OSHAPPython
★ Featured
Forecasting & Research

Rossmann Store Sales Forecasting

  • Built a LightGBM forecasting pipeline for 1,115 Rossmann stores across roughly 1.02M historical rows.
  • Used expanding-window validation with train-only store aggregates to keep the evaluation leakage-safe.
  • Reached 11.6% mean RMSPE against a 24.7% store-weekday median baseline; these are local validation results, not a Kaggle leaderboard score.
LightGBMOptunaPythonstatsforecast
ML & NLP

Two-Tower Movie Recommender

  • Built a two-stage recommender over 25M MovieLens ratings, using PyTorch Two-Tower retrieval with FAISS followed by LightGBM ranking.
  • Diagnosed embedding collapse (cosine similarity to the centroid measuring 1.000) and fixed it by swapping BatchNorm for LayerNorm, taking Recall@10 from 0.002 to a final 0.147.
  • Evaluated the engagement/diversity trade-off offline through a Pareto frontier and SNIPS counterfactual evaluation.
PyTorchFAISSLightGBMPython
Data Engineering

Lead Conversion Data Product

  • Designed a relational schema for lead-to-member conversion and containerised the Postgres database with Docker.
  • Built a dashboard prototype so business managers can explore 4 revenue KPIs directly.
PostgreSQLDockerPythonSQL
ML & NLP

NER in the CSIRO Adverse Drug Event Corpus

  • Fine-tuned BioBERT to extract adverse drug reactions from noisy biomedical text on the CSIRO CADEC corpus.
  • Used Focal Loss to counter severe class imbalance between common and rare entity types.
  • Lifted rare-entity F1 by over 20 points while reaching 88.92% overall accuracy (F1 88.04%).
BioBERTPyTorchHugging FaceFocal Loss
Data Engineering

Fashion Analyzer

  • Built PySpark pipelines over 2 years of fashion sales data spanning 9 categories, 10 styles and 5 regions.
  • Aggregated by category, style and region to surface assortment signals rather than raw transaction counts.
  • Found mid-range pricing ($50 to $150) drives the highest volume, with style preferences splitting sharply by market.
PySparkPythonData Viz
Forecasting & Research

Energy Time-Series Forecasting for Hydropower

  • Forecast hydro generation across Eastern India from 16 months of National Power Portal data.
  • Selected between ARIMA, SARIMA and Prophet using AIC and BIC rather than picking a model by eye.
  • Best model reached an RMSE of 153.55 and MAE of 78.03 on held-out months.
SARIMAARIMAProphetPython
Forecasting & Research

Reinforcement Learning Benchmarks

  • Implemented REINFORCE, Actor-Critic and A2C policy-gradient algorithms on CartPole-v1 in TensorFlow.
  • Extended the environment with transaction-cost penalties to simulate real-world P&L drift.
  • Ran 5 independent 1M-step experiments per method and compared final reward stability, not just peak reward.
TensorFlowPythonOpenAI Gym
Forecasting & Research

Graph Anonymization: Topology Preservation vs Privacy

  • Benchmarked 4 anonymization methods across 5 real-world graphs for structural preservation and re-identification risk.
  • Measured community-structure degradation separately from degree-distribution degradation, rather than a single blended score.
  • Found modularity shift stayed under 3% and re-identification risk under 1%, but aggressive anonymization degrades community structure faster than degree distribution.
PythonNetworkX
Data Engineering

where-the-money-goes-next

  • Built a fraud-detection pipeline scoring accounts most likely receiving money from authorised-payment scams, across 6.9M synthetic transactions.
  • Benchmarked a hand-written rule against logistic regression and XGBoost rather than assuming the ML model would win.
  • Set the alert threshold to a fixed daily analyst-review budget, with 173 tests gating the pipeline including a temporal-leakage check.
PythonpandasNumPyscikit-learnXGBoostSciPy
Forecasting & Research

Bayesian Media Mix Model for Advertising Spend

  • Built a Bayesian media mix model estimating each advertising channel's contribution to sales from two years of weekly data.
  • Ran a controlled comparison isolating spend burstiness as the reason the saturation knee goes unlearned on most channels.
  • Validated the PyMC inference with simulation-based calibration (15 of 15 parameters passed) and 139 automated tests.
PythonPyMCPyTensorArviZNumPypandas