GeoHab
Mapping the seafloor at Refuge Cove. A geospatial ML project using bathymetry, backscatter and labelled training points to predict five benthic habitat classes.

Inside the project Models, results & deployment decisions
The problem
Predict the habitat at unseen coordinates from underwater multibeam data. The five classes are imbalanced, so evaluation uses support-weighted F1.
My approach
Extract raster values, engineer spatial grids and compare LightGBM experiments. The two-stage stack combines a rich first-stage meta-model with a seed-averaged second-stage classifier.
What mattered
Spatial scale and stable features mattered more than complexity. Terrain, texture and clustering features did not consistently generalize across the leaderboard splits.
Choosing what to ship
The public leaderboard used 31% of the test data; the private leaderboard used 69%. I kept the public and private results separate when choosing models for the interactive app.
| Model / decision | Private | Public |
|---|---|---|
| Buffered Bayes adjusted ensembleDefault · stable across splits, fast feature pipeline | 0.84295 | 0.84950 |
| Meta-stacked modelAlternative · best private score, heavier feature computation | 0.85875 | 0.79777 |
| CNN + LGBM blendNot deployed · public/private gap suggested overfitting to the public split | 0.84447 | 0.91568 |
The generalization lesson
A first-place public result and a thirteenth-place private result tell different stories. Spatial distribution shifts exposed the limits of standard cross-validation. Spatial block validation and better handling of those shifts are the next steps.
The CNN experiments incorporate publicly shared out-of-fold predictions by Matteo. Marine mapping data: Deakin Marine Mapping Group, CC BY 4.0.



