Hult Business Challenge II · Team 4 · 2026 · I led the modelling
Urban heat islands from satellite data
Classifying urban heat-island intensity (Low, Medium, High) at 100 m resolution from Sentinel-2, Landsat-8 and elevation data. The models learn on Rio de Janeiro and Santiago, then transfer to Freetown, Sierra Leone, a city with no labels of its own.
My role
The problem
Cities trap heat unevenly, and the hottest blocks are where heat stress hits hardest. Ground sensors are sparse, so the question is whether satellite data can map heat-island intensity, and whether a model trained in one city still works in another.
Data and cloud setup
| Imagery | Sentinel-2 L2A, Landsat-8 thermal, a digital elevation model and 3D building footprints, from Microsoft Planetary Computer |
|---|---|
| Points | 50,150 labelled points across Rio and Santiago; Freetown unlabelled |
| Features | 25 spectral indices, land-surface temperature, elevation, building morphology and 6 interaction terms |
| Compute | Microsoft Fabric notebooks reading and writing the Lakehouse; scenes loaded lazily in 2048 × 2048 chunks as uint16 and released after sampling; intermediate results cached as Parquet |
Approach and results
Rio: tune for a strong heat signal. A class-balanced XGBoost beat Random Forest in 5-fold cross-validation (0.959 against 0.947). Thermal bands dominate, then moisture (NDMI) and land-surface temperature, then building compactness.
Santiago: terrain, not materials. Half the labels are Medium and overlap both extremes. Elevation and its interaction with temperature are the top features, because the Andean basin creates gradients surface materials don't explain. A depth-limited Random Forest on quantile-transformed features handled the ambiguous class better than boosting.
Freetown: the transfer problem. Raw temperatures don't carry across cities: Santiago's "High" sits near Rio's "Low".
So the transfer model:
- Quantile-transforms five thermal and spectral features separately for each city, so "hot for Freetown" lines up with "hot for Rio".
- Compresses them to 3 principal components, which keep 96% of the variance.
- Trains a one-vs-rest specialist per class on the source city that matches Freetown best for that class.
- Averages the Random Forest and XGBoost specialists.
This scored 0.58 on the challenge leaderboard's hidden labels.
Limitations
- Freetown has no ground truth, so its error can't be broken down by class.
- Each city is one satellite composite from one date window. Seasonal data is the natural next step.
Reproduce it
git clone https://github.com/youness-yach/uhi-business-challenge.git
cd uhi-business-challenge && pip install -e .
jupyter nbconvert --to notebook --execute notebooks/02_uhi_classification.ipynb
The Rio and Santiago results reproduce in about 4 minutes on a laptop. The same notebook runs on Microsoft Fabric.