A collection of things I’ve made

Projects.

Intelligent systems. A little room for play.

ML & systems

04Web & interaction
01 / GeoHab MLWG · 2026 Geospatial machine learning

GeoHab

Mapping the seafloor at Refuge Cove. A geospatial ML project using bathymetry, backscatter and labelled training points to predict five benthic habitat classes.

Refuge Cove / survey map
Backscatter map of Refuge Cove with 6,256 ground-truth training locations coloured by habitat class, over bathymetric contour lines
0.85875Best private weighted F1 · meta-stackWeighted F1 runs 0–1; 1 is perfect
0.84295Shipped model · private weighted F1Chosen for holding steady across both splits
1132652#1#13PublicPrivate1st on the public board, 13th of 52 on the private oneThe public split was 31% of the test data
Inside the project Models, results & deployment decisions

The problem

Predict the habitat at unseen coordinates from underwater multibeam data. The five classes are imbalanced, so evaluation uses support-weighted F1.

My approach

Extract raster values, engineer spatial grids and compare LightGBM experiments. The two-stage stack combines a rich first-stage meta-model with a seed-averaged second-stage classifier.

What mattered

Spatial scale and stable features mattered more than complexity. Terrain, texture and clustering features did not consistently generalize across the leaderboard splits.

Choosing what to ship

The public leaderboard used 31% of the test data; the private leaderboard used 69%. I kept the public and private results separate when choosing models for the interactive app.

GeoHab model comparison · weighted F1
Model / decisionPrivatePublic
Buffered Bayes adjusted ensembleDefault · stable across splits, fast feature pipeline0.842950.84950
Meta-stacked modelAlternative · best private score, heavier feature computation0.858750.79777
CNN + LGBM blendNot deployed · public/private gap suggested overfitting to the public split0.844470.91568

The generalization lesson

A first-place public result and a thirteenth-place private result tell different stories. Spatial distribution shifts exposed the limits of standard cross-validation. Spatial block validation and better handling of those shifts are the next steps.

The CNN experiments incorporate publicly shared out-of-fold predictions by Matteo. Marine mapping data: Deakin Marine Mapping Group, CC BY 4.0.

02 / Kaggle Playground · 2026 Selected project

Predicting F1 Pit Stops

A lap-level machine-learning system that turns tyre, timing, position and race context into an actionable pit-stop probability.

ROC AUC / experiment progression
0.95369Best saved blend · private ROC AUC
+0.01760Private AUC over the LGBM baseline
Inside the project Experiments & prediction service

From 439,140 lap-level examples to a next-lap pit probability. Feature engineering and an out-of-fold LightGBM + RealMLP blend improved the private score from 0.93609 to 0.95369.

Finding the strategy signal

I explored tyre life, stint, lap timing and race context in a synthetic competition dataset with roughly an 80:20 non-pit/pit balance. Race-level variation and anomalous 2023 pit rates made validation an important part of the investigation.

From baseline to blend

The experiments cover LightGBM, XGBoost and RealMLP. The best saved blend combines 42.5% LightGBM and 57.5% RealMLP, with 0.953876 out-of-fold AUC and 0.95328 public AUC. The private split contains 80% of the test data; the public split contains 20%.

Validation changed between experiments: the baseline used a year holdout, one XGBoost run grouped by race, and later feature-engineered runs used stratified folds. Their local scores are not directly comparable under one fixed validation protocol.

Taking the model into an app

The live app serves a native LightGBM model through FastAPI, separately from the best competition blend. Startup checks verify the model, metadata checksum, feature contract and exported smoke test. The model loads once and inference uses one CPU thread.

Competition: Kaggle Playground Series S6E5. ROC AUC measures ranking quality; these synthetic-data results do not establish performance in live Formula 1 races.

03 / Desktop app · 2026 Open source

VectorFlow

A local Windows desktop app that automatically redraws a video as animated cubic Bézier curves. Choose a video, convert it, then preview or export the result. No manual tracing.

P₀P₁P₂P₃
6Export formats · MP4, animated SVG, SVG & PNG sequences, JSON, ProRes 4444
0Uploads · every frame is processed on your own machine
Inside the project Tracing, tracking & editing

Every frame becomes real cubic curves: the exported SVGs contain actual C path commands, and each project stores the control points, stable path IDs and motion per frame.

How it works

OpenCV decodes frames at the chosen output rate and smooths nearly stationary pixels. Optical flow and conservative shape matching give each path a persistent ID, so confident fits reuse their curve topology from frame to frame. Filled-colour mode learns a palette across the clip, then fits closed contours and holes.

Editing after conversion

A cleanup workspace lets you drag a curve's four handles, delete or simplify paths, and apply edits to one frame or a whole tracked segment. Text and artwork can be replaced inside a marked region and follow its motion. Edits are reversible and never overwrite the original frame data.

Built to stay responsive

The PySide6 interface keeps conversion and FFmpeg export off the UI thread and writes frames to disk incrementally instead of holding the whole video in memory.

Tracking is geometric rather than semantic, so fast deformation, occlusion or hard cuts can break a track. Those limits are documented in the repository.

Embedding match / system overview
04 / Product engineering Internship project

Folio

A job portal for design students. I designed the user flow and data flow for the whole system and built its matchmaking pipeline, which ranks students against roles using vector embeddings stored in ChromaDB.

FocusEmbeddings · ChromaDB · System design
Since handed over to another team
Competition results & notebooksThe experiments behind the work
Kaggle / the competition archive

Hall of fame

Experiments put to the test.
Notebooks shared along the way.

31Competitions entered
07Bronze medal notebooks
1084/61079Kaggle Notebooks Expert rank
Kaggle ↗

Competition results

Best results captured on Kaggle
  1. 01

    AI Agent Security – Multi-Step Tool Attacks

    Featured · Code competition

    457 / 4,186Rank / teams
    Top 11%
  2. 02

    Predicting Heart Disease

    Playground Series · S6E2

    604 / 4,370Rank / teams
    Top 14%
  3. 03

    Predict Customer Churn

    Playground Series · S6E3

    742 / 4,142Rank / teams
    Top 18%
  4. 04

    Cyber-Physical Anomaly Detection for DER Systems

    Community · DER security

    Same score as 3rd place; ranked 11th because the submission came later.

    11 / 61Rank / teams
    Top 19%
  5. 05

    Predicting F1 Pit Stops

    Playground Series · S6E5

    615 / 3,022Rank / teams
    Top 21%
  6. 06

    Predicting Stellar Class

    Playground Series · S6E6

    600 / 2,816Rank / teams
    Top 22%
  7. 07

    GeoHab 2026 MLWG Competition

    Community · Benthic habitat classification

    1st on the public leaderboard · 13th on the private leaderboard.

    13 / 52Rank / teams
    Top 25%

Notebook shelf

Community recognition
Notebook / 01

Irrigation Needs | Detailed EDA

Notebook previewirrigation-needs.ipynb

Ranks, medals and votes reflect the supplied Kaggle snapshots; leaderboard notes are provided by the author. Results are ordered by rank / teams, with top percentages rounded up.

Web & interaction

08