WELCOME
Devika Rajasekar
Data & Analytics Engineer
Netherlands · She/her

“Why?” is my favorite data point.

I work on data infrastructure and machine learning models. Also, I will try any gelato flavor once, and I bring the same energy to a new tool or dataset. I learn it properly, get obsessed, build something with it. That’s basically the story behind every project below.

Devika Rajasekar professional data nerd, part-time ice cream critic
⠀⠀⠀⠀⢠⡶⠚⢷⣤⡀⠀⠀⠀⠀⠀⣲⡶⠛⠻⣆⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠀⠀⢠⡿⠁⠀⠀⠙⣷⣄⠀⢀⣴⡟⠁⠀⠀⢷⢹⡆⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠀⠀⣾⠃⠀⠠⠶⠚⠛⠛⠛⠛⠋⠀⠀⣀⡀⢸⠈⣿⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠀⢸⣏⡔⠋⠀⠀⠀⠀⠀⠀⠀⠀⠀⠚⠉⠉⣿⠀⢹⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠀⢾⠏⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠸⠀⢸⡇⠀⠀⠀⠀⠀⠀⠀⠀
⠀⢠⣿⢠⣶⡆⠀⠀⠀⠀⣀⣀⠀⠀⠀⠀⠀⠀⠀⠀⢸⡇⠀⠀⠀⠀⠀⠀⠀⠀
⢒⡾⠁⠘⠟⠁⠀⠀⠀⠀⣿⣿⡆⠀⠀⠀⠀⠀⠀⠀⢸⡇⠀⠀⠀⠀⠀⠀⠀⠀
⠉⣧⠀⠀⠀⠀⠃⠀⠀⠀⠈⠉⠠⣍⠀⠀⠀⠀⠀⠀⣸⡇⢀⣤⠶⠛⠛⠻⢦⣄
⠀⠸⣧⡀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⣰⡟⣴⠟⠁⠀⠀⠀⠀⠀⢻
⠀⠀⠀⠛⣷⡦⠀⠀⠀⠀⠀⠀⠀⠀⣀⣀⣤⡴⠞⠋⢠⡟⠀⠀⠀⠀⠀⠀⢀⡾
⠀⠀⠀⢰⡿⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠉⠳⣤⡀⢸⠃⠀⠀⠀⠀⢠⡶⠟⠁
⠀⠀⠀⣸⠇⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠘⢷⣹⡄⠀⠀⠀⠀⣼⠀⠀⠀
⠀⠀⠀⣿⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠈⢿⣇⠀⠀⠀⠀⢹⡄⠀⠀
⠀⠀⠀⢸⡀⢀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠈⣿⡄⠀⠀⠀⠈⣧⠀⠀
⠀⠀⠀⢸⡇⠘⡇⠀⠀⠀⠀⠀⠀⠀⣀⠀⠀⠀⠀⠀⠀⢸⣿⠀⠀⠀⠀⢹⡇⠀
⠀⠀⠀⢸⡇⠀⠙⠀⠀⠀⠀⠀⢠⠞⠁⠀⠀⠀⠀⠀⠀⠀⣿⠇⠀⠀⠀⢸⡇⠀
⠀⠀⠀⢸⡇⠀⢸⡆⠀⠀⠀⠀⣟⠀⠀⠀⠀⠀⠀⠀⠀⠀⠛⠀⠀⠀⠀⣸⠇⠀
⠀⠀⠀⢸⣿⠀⠀⡇⠀⠀⠀⠀⣿⡀⠀⠀⠀⠀⠀⠀⠀⢀⡇⠀⠀⢀⣴⡟⠁⠀
⠀⠀⠀⠘⠿⠶⢶⢧⣦⣦⡴⢾⣥⣽⣤⣤⣤⣤⣤⣤⡴⣯⡤⠴⠶⠛⠋⠀⠀⠀
ps. drink some water
psst, the stickers moveTRY DRAGGING ME
Meet the engineer

Hello,
I’m Devika!

I’m wrapping up an MSc in Computer Science (Data Science) at Leiden, and I work with data at Prysmian in Delft. I like being in all the parts of a data problem: shaping raw tables, building models, judging when their output can be trusted, and getting a result into a form people can act on. I like different parts on different days, and I’m growing most into the engineering that holds it all together.

The human side of things has always pulled at me, sociology, anthropology, biology, so taking my CS degree alongside marketing and lean-startup courses felt natural. Who reads the output, and what do they change because of it? It’s the piece I look for first in any new project.

off the clock, I’m all about

cinema ticket, movies & shows coffee stamp spicy food stamp
Devika eating gelato
best gelato I’ve had in my life
Where I’ve worked

Experience

2025AUG – NOW

HR Digital Data Science Intern

Prysmian · Netherlands
  • Replaced monthly manual KPI consolidation across 53 plant-level Excel sheets with a single live Qlik Sense dashboard fed from a central data lake, covering 15 countries and 4 regions (CEE, NE, SE, UK), trusted by senior stakeholders.
  • Translated cross-functional reporting requirements into reusable data models, KPI definitions and dashboard logic covering variance analysis and conditional indicators for senior leadership.
  • Running my MSc thesis on AutoML explainability: whether SHAP feature-importance stays consistent across the Rashomon set of near-optimal models, with AutoGluon & H2O experiments deployed on an HPC cluster.
Qlik SensePower AutomateAutoMLSHAPHPC
2023MAY – JUL

Deep Learning & Image Processing Research Intern

CCPS, VIT Chennai · India
  • Built a custom U-Net variant (AdaptUNet) to detect colorectal polyps from 2,000+ medical images: 0.9104 Dice, 0.9880 balanced accuracy.
  • Ran cross-dataset generalisation and data-validation checks. Published as first-author in Elsevier Heliyon (2024) ↗.
PyTorchU-NetMedical Imaging
2023JUL – NOV

Data Analysis Intern

School of CSE, VIT Chennai · India
  • Designed a deep neural network in R for pulsar-star classification: 98% precision, 0.8749 Kappa on the HTRU2 dataset.
  • Led reproducibility & compliance checks to IEEE academic QA standards, underpinning a successful conference publication ↗.
RDeep LearningIEEE
2022AUG – DEC

Research Assistant

School of CSE, VIT Chennai · India
ResearchRoboticsLiterature Review
whimsical illustrated cat covered in colourful bows
A pinboard of projects

Some of my work

Data Engineering

paytrail: Payments Medallion Lakehouse

A governed bronze/silver/gold lakehouse over 6.3M synthetic payments on Databricks and Delta: idempotent out-of-order-safe loads, Unity Catalog governance, 51 CI-gated dbt tests, and a code-first BI dashboard.

6.3Msynthetic payments, penny-reconciled
39% fasterheavy rollup after OPTIMIZE + Z-order
DatabricksDelta LakeUnity CatalogdbtAzureGitHub Actions
Data Engineering

Workplace Safety Analytics Pipeline

Built an end-to-end batch pipeline ingesting 688K+ OSHA workplace-injury records into a PostgreSQL star-schema warehouse (1 fact + 5 dimensions) with dbt, orchestrated by a 7-task Airflow DAG, runnable with a single docker compose up. Added an LLM enrichment layer using a locally-hosted model via Ollama that classified incident narratives at 96.5% coverage at $0 inference cost. Enforced data quality with 81 automated tests wired into a GitHub Actions CI pipeline that fails the build on any data-contract violation.

688K+OSHA records modelled
81automated tests, green in CI
Apache AirflowdbtPostgreSQLDockerGitHub ActionsOllamaPython
Forecasting & Research

Rossmann Store Sales Forecasting

Forecast six weeks of sales for 1,115 stores with a single LightGBM model, placing top 5% of 3,738 Kaggle teams. Used walk-forward cross-validation to prevent lookahead leakage. A GitHub Actions workflow runs automated tests on the feature engineering pipeline on every push, keeping model inputs validated and the project reproducible.

top 5%of 3,738 Kaggle teams
11.5%RMSPE · walk-forward CV, no leakage
LightGBMOptunaPythonstatsforecast
ML & NLP

Two-Tower Movie Recommender

A two-stage movie recommender on MovieLens 25M: Two-Tower retrieval (PyTorch, FAISS) feeding a 306-feature LightGBM ranker that scores Recall@10 of 0.147 and 0.9206 val AUC. The engagement/diversity trade-off is simulated offline as a 12-point Pareto frontier, with candidate policies validated by SNIPS counterfactual evaluation.

25M ratingsMovieLens, 13,176-item catalogue
12-point Paretoengagement/diversity, SNIPS-validated
PyTorchFAISSLightGBMPython
Forecasting & Research

SHAP Stability Across the Rashomon Set

Does a model's explanation survive being replaced by an equally accurate one? My thesis tests this across the Rashomon set on 6 forecasting benchmarks.

6 benchmarksincl. ETT & M4 Monthly
AutoGluon + H2Odeployed on Leiden's ALICE HPC
AutoGluonH2OSHAPPython
ML & NLP

The Edit: H&M Recommender

A two-stage recommender on 31M H&M transactions: BigQuery SQL retrieval feeding a CatBoost ranker (MAP@12 0.029 vs 0.0053 baseline, a 5.5x lift with a 95% bootstrap CI), plus a temporal-leakage gate, diversity guardrails, a typed FastAPI endpoint, and 80+ tests.

5.5×MAP@12 lift over popularity baseline
31MH&M transactions, built in BigQuery free tier
BigQueryCatBoostFastAPIPythonSQL
Data Engineering

Multi-System Data Cleaning & Entity Reconciliation

Consolidated fragmented payment, CRM and ERP data from Stripe, Salesforce and NetSuite into a single customer view across ~2M rows, resolving 98% of foreign-key mismatches.

98%FK mismatches resolved
~2Mrows reconciled
Entity ResolutionSQLER DiagramsRule-based Matching
Data Engineering

Lead Conversion Data Product

Designed a relational schema for lead-to-member conversion, containerised the Postgres database with Docker, and built a dashboard prototype letting business managers explore 4 revenue KPIs directly.

Postgrescontainerised w/ Docker
KPIdashboard prototype
PostgreSQLDockerPythonSQL
ML & NLP

NER in the CSIRO Adverse Drug Event Corpus

Fine-tuned BioBERT to extract adverse drug reactions from noisy biomedical text, using Focal Loss to handle severe class imbalance. Rare-entity classes improved by over 20% F1, with 88.92% overall accuracy.

88.92%accuracy (F1 88.04%)
+20%rare-entity F1 lift
BioBERTPyTorchHugging FaceFocal Loss
Data Engineering

Fashion Analyzer

Built PySpark pipelines across 2 years of fashion sales data covering 9 categories, 10 styles and 5 regions. The pipeline found that mid-range pricing ($50 to $150) drives the highest volume and that style preferences split sharply by market, framed as category-level assortment signals.

PySparkdistributed pipelines
CSV / JSONcurated insight exports
PySparkPythonData Viz
Forecasting & Research

Energy Time-Series Forecasting for Hydropower

Forecast hydro generation across Eastern India from 16 months of National Power Portal data, with model selection driven by AIC and BIC.

153.55RMSE
78.03MAE
SARIMAARIMAProphetPython
Forecasting & Research

Reinforcement Learning Benchmarks

Implemented REINFORCE, Actor-Critic, and A2C policy gradient algorithms on CartPole-v1 in TensorFlow, then ran 5 independent 1M-step experiments with logging, smoothing, and comparison plots. Extended the environment with transaction-cost penalties to simulate real-world P&L drift and evaluated final reward stability across methods.

1M-step experiments on CartPole-v1
TensorFlowPythonOpenAI Gym
Forecasting & Research

Graph Anonymization: Topology Preservation vs Privacy

Evaluated 4 anonymization methods across 5 real-world graphs, measuring structural preservation (modularity shift under 3%) and re-identification risk (under 1%). Findings show that aggressive anonymization degrades community structure faster than degree distribution, with implications for privacy-utility trade-offs in graph-based ML pipelines.

4anonymization methods benchmarked
5real-world graphs, modularity shift <3%
PythonNetworkX
My toolkit

Tools & skills

Languages

PythonRSQL

ML & Deep Learning

PyTorchTensorFlowScikit-LearnTransformersComputer VisionPySparkReinforcement Learning

Data Engineering

DatabricksDelta LakeUnity CatalogdbtAirflowDatabricks Asset BundlesDockerGitUiPath

Analysis & Viz

PandasNumPyMatplotlibSeabornPower BITableauQlik SenseExcel

Cloud

Microsoft Azure (ADLS Gen2)AWS (S3 · EC2 · Lambda)Google Cloud

Ways of working

Agile / ScrumData ValidationRequirement GatheringUser Story CreationStakeholder CommsAutomated Testing
On paper

Background

MSc Computer Science: Data Science
Leiden University, Netherlands
SEP 2024 – AUG 2026 (expected)
Coursework: Automated Machine Learning · Recommender Systems · Reinforcement Learning · Social Network Analysis · Text Mining · Data Mining. Thesis: whether SHAP explanations stay consistent when you pick from a set of near-optimal models rather than a single best model.
BTech Computer Science Engineering: AI & Robotics
Vellore Institute of Technology (VIT), India
SEP 2020 – JUN 2024 · CGPA 8.5 / 10
Machine Learning · Linear Algebra · Deep Learning · Data Analysis, taken alongside sociology, marketing & lean-startup.

Publications

Elsevier Heliyon · 2024 · first author

0.9104 Dice · cross-dataset generalisation tests

ICoFT MADE 2022 · 2022 · co-author

review of 20+ papers on robotics in search & rescue

Languages

EnglishNATIVE
Hindi & TamilFLUENT
DutchA2 · BASIC

Leadership & engagement

General Secretary · CodeChef VIT Chennai
JAN – JUL 2023

Drove recruitment to 500+ applicants. Ran coding events, workshops and seminars to grow a vibrant coding culture.

Active Member · IEEE Women in Engineering, VIT
MAY – DEC 2021

Advocated for women in STEM and organised professional-development events across diverse stakeholder groups.

Licenses & certifications

View all on LinkedIn ↗

AWS: Introduction to Machine Learning on AWS · Google Cloud: Big Data & Machine Learning Fundamentals · IBM: Artificial Intelligence Analyst