★ Featured Data Engineering
paytrail: Payments Medallion Lakehouse
- Built a Databricks medallion lakehouse taking 6.3M synthetic payments from Azure storage through bronze, silver and gold layers.
- Designed idempotent incremental loads for late and out-of-order data, with Unity Catalog governance and 51 dbt tests gated in CI.
- Reconciled settlement outputs to source at penny level and reduced the heaviest rollup runtime by 39% after optimisation.
DatabricksDelta LakeUnity CatalogdbtAzureGitHub Actions
★ Featured Data Engineering
Cross-System Customer Data Reconciliation
- Reconciled customer records across Salesforce, NetSuite and Stripe-style data, starting from 12,566 supplied case rows.
- Built deterministic exact-key cleaning for duplicate records, missing identifiers, mixed date formats and European/US currency formats.
- Kept unresolved links visible instead of inferring them: 49 NetSuite rows missing a Stripe ID and 675 unmatched Stripe transactions surfaced as exceptions.
PythonpandasDeterministic Matching
★ Featured Data Engineering
Workplace Safety Analytics Pipeline
- Built a batch pipeline ingesting 688K+ OSHA injury records into a PostgreSQL star schema, orchestrated through a seven-task Airflow DAG.
- Added local LLM enrichment for free-text incident narratives using Ollama, keeping the pipeline reproducible and inference cost-free.
- Enforced 81 automated data and contract tests through dbt and GitHub Actions, with failed checks blocking the build.
Apache AirflowdbtPostgreSQLDockerGitHub ActionsOllamaPython
★ Featured ML & NLP
The Edit: H&M Recommender
- Built a two-stage recommender over 31M H&M transactions, using BigQuery SQL retrieval followed by a CatBoost ranking model.
- Improved MAP@12 from 0.0053 to 0.0292, with a 95% bootstrap confidence interval confirming the lift was stable.
- Added a temporal-leakage gate, diversity checks, 80+ automated tests, and a typed FastAPI prediction endpoint.
BigQueryCatBoostFastAPIPythonSQL
★ Featured Forecasting & Research
SHAP Stability Across the Rashomon Set
- Tested whether SHAP feature rankings remain stable when a forecasting model is replaced by another near-optimal model.
- Ran AutoGluon and H2O experiments across six forecasting benchmarks, multiple temporal splits and Rashomon thresholds on Leiden's ALICE HPC.
- Found H2O built larger, more varied Rashomon sets than AutoGluon (up to 15 models across 5 families, against up to 5 in 1), yet both frameworks' stability agreed within 0.04 Spearman rho on every dataset (0.77 to 0.98 overall).
AutoGluonH2OSHAPPython
★ Featured Forecasting & Research
Rossmann Store Sales Forecasting
- Built a LightGBM forecasting pipeline for 1,115 Rossmann stores across roughly 1.02M historical rows.
- Used expanding-window validation with train-only store aggregates to keep the evaluation leakage-safe.
- Reached 11.6% mean RMSPE against a 24.7% store-weekday median baseline; these are local validation results, not a Kaggle leaderboard score.
LightGBMOptunaPythonstatsforecast
ML & NLP
Two-Tower Movie Recommender
- Built a two-stage recommender over 25M MovieLens ratings, using PyTorch Two-Tower retrieval with FAISS followed by LightGBM ranking.
- Diagnosed embedding collapse (cosine similarity to the centroid measuring 1.000) and fixed it by swapping BatchNorm for LayerNorm, taking Recall@10 from 0.002 to a final 0.147.
- Evaluated the engagement/diversity trade-off offline through a Pareto frontier and SNIPS counterfactual evaluation.
PyTorchFAISSLightGBMPython
Data Engineering
Lead Conversion Data Product
- Designed a relational schema for lead-to-member conversion and containerised the Postgres database with Docker.
- Built a dashboard prototype so business managers can explore 4 revenue KPIs directly.
PostgreSQLDockerPythonSQL
ML & NLP
NER in the CSIRO Adverse Drug Event Corpus
- Fine-tuned BioBERT to extract adverse drug reactions from noisy biomedical text on the CSIRO CADEC corpus.
- Used Focal Loss to counter severe class imbalance between common and rare entity types.
- Lifted rare-entity F1 by over 20 points while reaching 88.92% overall accuracy (F1 88.04%).
BioBERTPyTorchHugging FaceFocal Loss
Data Engineering
Fashion Analyzer
- Built PySpark pipelines over 2 years of fashion sales data spanning 9 categories, 10 styles and 5 regions.
- Aggregated by category, style and region to surface assortment signals rather than raw transaction counts.
- Found mid-range pricing ($50 to $150) drives the highest volume, with style preferences splitting sharply by market.
PySparkPythonData Viz
Forecasting & Research
Energy Time-Series Forecasting for Hydropower
- Forecast hydro generation across Eastern India from 16 months of National Power Portal data.
- Selected between ARIMA, SARIMA and Prophet using AIC and BIC rather than picking a model by eye.
- Best model reached an RMSE of 153.55 and MAE of 78.03 on held-out months.
SARIMAARIMAProphetPython
Forecasting & Research
Reinforcement Learning Benchmarks
- Implemented REINFORCE, Actor-Critic and A2C policy-gradient algorithms on CartPole-v1 in TensorFlow.
- Extended the environment with transaction-cost penalties to simulate real-world P&L drift.
- Ran 5 independent 1M-step experiments per method and compared final reward stability, not just peak reward.
TensorFlowPythonOpenAI Gym
Forecasting & Research
Graph Anonymization: Topology Preservation vs Privacy
- Benchmarked 4 anonymization methods across 5 real-world graphs for structural preservation and re-identification risk.
- Measured community-structure degradation separately from degree-distribution degradation, rather than a single blended score.
- Found modularity shift stayed under 3% and re-identification risk under 1%, but aggressive anonymization degrades community structure faster than degree distribution.
PythonNetworkX
Data Engineering
where-the-money-goes-next
- Built a fraud-detection pipeline scoring accounts most likely receiving money from authorised-payment scams, across 6.9M synthetic transactions.
- Benchmarked a hand-written rule against logistic regression and XGBoost rather than assuming the ML model would win.
- Set the alert threshold to a fixed daily analyst-review budget, with 173 tests gating the pipeline including a temporal-leakage check.
PythonpandasNumPyscikit-learnXGBoostSciPy
Forecasting & Research
Bayesian Media Mix Model for Advertising Spend
- Built a Bayesian media mix model estimating each advertising channel's contribution to sales from two years of weekly data.
- Ran a controlled comparison isolating spend burstiness as the reason the saturation knee goes unlearned on most channels.
- Validated the PyMC inference with simulation-based calibration (15 of 15 parameters passed) and 139 automated tests.
PythonPyMCPyTensorArviZNumPypandas