← All projects

Hult Business Challenge II · Team 4 · 2026 · I led the modelling

Urban heat islands from satellite data

Classifying urban heat-island intensity (Low, Medium, High) at 100 m resolution from Sentinel-2, Landsat-8 and elevation data. The models learn on Rio de Janeiro and Santiago, then transfer to Freetown, Sierra Leone, a city with no labels of its own.

MethodsFeature engineering (spectral indices)Random ForestXGBoostClass balancingStratified cross-validationQuantile transform + PCAOne-vs-rest transfer ToolsPythonscikit-learnXGBoostxarrayrasteriogeopandasMicrosoft FabricPlanetary Computer DomainGeospatial machine learning · climate risk · cross-domain transfer
0.959
weighted F1, Rio de Janeiro (XGBoost, 5-fold CV)
0.690
weighted F1, Santiago (Random Forest, 20% hold-out)
0.58
weighted F1, Freetown, a city the model never saw labels for (challenge leaderboard)

My role

I ran the machine-learning work: the model search and selection for Rio and Santiago (model families, resolutions, tuning, and picking the best fit for each city), and the transfer model that predicts Freetown. Satellite extraction was built by teammate Mickias Ambaye, whose uhi_pipe package is credited in the repo.

The problem

Cities trap heat unevenly, and the hottest blocks are where heat stress hits hardest. Ground sensors are sparse, so the question is whether satellite data can map heat-island intensity, and whether a model trained in one city still works in another.

Data and cloud setup

ImagerySentinel-2 L2A, Landsat-8 thermal, a digital elevation model and 3D building footprints, from Microsoft Planetary Computer
Points50,150 labelled points across Rio and Santiago; Freetown unlabelled
Features25 spectral indices, land-surface temperature, elevation, building morphology and 6 interaction terms
ComputeMicrosoft Fabric notebooks reading and writing the Lakehouse; scenes loaded lazily in 2048 × 2048 chunks as uint16 and released after sampling; intermediate results cached as Parquet

Approach and results

Rio: tune for a strong heat signal. A class-balanced XGBoost beat Random Forest in 5-fold cross-validation (0.959 against 0.947). Thermal bands dominate, then moisture (NDMI) and land-surface temperature, then building compactness.

Feature importance for the Rio model, led by thermal bands
Figure 1. What drives the Rio model: thermal infrared first, then moisture and surface temperature. Source

Santiago: terrain, not materials. Half the labels are Medium and overlap both extremes. Elevation and its interaction with temperature are the top features, because the Andean basin creates gradients surface materials don't explain. A depth-limited Random Forest on quantile-transformed features handled the ambiguous class better than boosting.

Freetown: the transfer problem. Raw temperatures don't carry across cities: Santiago's "High" sits near Rio's "Low".

Land-surface temperature by heat class in each city, showing the shift between cities
Figure 2. Land-surface temperature by class and city. A model trained on Rio scores only 0.36 on Santiago, so absolute temperatures can't be transferred. Source

So the transfer model:

  1. Quantile-transforms five thermal and spectral features separately for each city, so "hot for Freetown" lines up with "hot for Rio".
  2. Compresses them to 3 principal components, which keep 96% of the variance.
  3. Trains a one-vs-rest specialist per class on the source city that matches Freetown best for that class.
  4. Averages the Random Forest and XGBoost specialists.

This scored 0.58 on the challenge leaderboard's hidden labels.

Heat-island class at every sample point in Rio, Santiago and Freetown
Figure 3. Heat-island class at every sample point: labelled in Rio and Santiago, predicted in Freetown. Source

Limitations

  • Freetown has no ground truth, so its error can't be broken down by class.
  • Each city is one satellite composite from one date window. Seasonal data is the natural next step.

Reproduce it

git clone https://github.com/youness-yach/uhi-business-challenge.git
cd uhi-business-challenge && pip install -e .
jupyter nbconvert --to notebook --execute notebooks/02_uhi_classification.ipynb

The Rio and Santiago results reproduce in about 4 minutes on a laptop. The same notebook runs on Microsoft Fabric.