Evaluating Sentinel-2 Super-Resolution for Geospatial Information Extraction
Super-resolution (SR) methods for Earth observation are commonly evaluated using reconstruction or perceptual image-quality metrics. For many geospatial applications, however, spatial sharpness alone is not sufficient. Index-based workflows depend on physically meaningful relationships between spectral bands, and spectral distortions introduced by SR can propagate directly into the resulting thematic products.
In our new publication in MDPI Geomatics, “Evaluating Sentinel-2 Super-Resolution for Geospatial Information Extraction: A Spectral and Thematic Assessment,” we evaluate this problem using two real hazard-mapping cases: the 2024 Valencia flood and the 2025 Palisades wildfire.
Experimental setup
We compare five model families for 4× super-resolution of the native 10 m Sentinel-2 RGB–NIR bands: LDSR-S2, OpenSR-SRGAN, SPAN, Mamba, and SWIN. Each model reconstructs bands B2, B3, B4, and B8 from 10 m to 2.5 m.
The resulting RGB–NIR products are then used within SEN2SR to reconstruct the six native 20 m red-edge and SWIR bands. This produces a ten-band Sentinel-2 product at a common spatial resolution of 2.5 m for every configuration.
All models are used without event-specific training or fine-tuning. The evaluation therefore measures how pretrained SR systems behave when applied to spectral conditions associated with flooding, burned vegetation, char, exposed soil, and other hazard-related surface changes.
Spectral consistency and spatial synthesis
We evaluate the SR products after aggregation back to the native Sentinel-2 grid using conventional reconstruction metrics and the OpenSR consistency and synthesis metrics. The results show a clear trade-off between preserving the original Sentinel-2 measurements and introducing additional high-frequency spatial content.
SWIN achieves the strongest native-grid reconstruction and spectral consistency among the learned models, but also has the lowest synthesis score. SRGAN shows the opposite behaviour: it introduces the most high-frequency content, but also produces larger radiometric and spectral deviations. SPAN, Mamba, and LDSR-S2 occupy different positions between these two extremes.
This result is important for downstream applications: a stronger spatial response does not necessarily imply a more faithful reconstruction of the underlying Sentinel-2 observation.
Flood and burn-scar mapping
We evaluate the reconstructed imagery using two physically interpretable spectral indices. Flood water is detected using MNDWI, based on the green and SWIR bands, while burn-scar mapping uses dNBR, derived from changes in NIR and SWIR reflectance between pre- and post-fire acquisitions.
For the Valencia flood, all learned configurations slightly improve the full-region agreement relative to bilinear interpolation. LDSR-S2 gives the highest full-region flood F1-score, increasing it from 0.085 to 0.091.
For the Palisades fire, SWIN provides the strongest learned-model full-region agreement, with an F1-score of 0.884 and an IoU of 0.792. The model ranking therefore depends strongly on the downstream task and evaluation criterion.
Most changes occur close to hazard boundaries
The largest effects of SR occur in the vicinity of mapped flood and fire boundaries. For flood mapping, the learned models increase edge-region detections by approximately 20–38%. Edge recall, F1-score, and IoU generally increase, but this is accompanied by a reduction in edge precision.
The same trade-off is visible in the fire case. SR can move mixed boundary pixels across the spectral-index decision threshold, increasing sensitivity to spatial structure near the event boundary while also introducing additional detections or boundary fragments.
Evaluating SR by fitness for use
No single model performs best across all criteria. SWIN provides the strongest native-grid spectral consistency and fire-task agreement, LDSR-S2 performs best for several flood and edge metrics, SPAN provides strong spatial consistency, Mamba achieves the lowest learned-model symmetric flood-boundary distance, and SRGAN produces the strongest high-frequency synthesis and edge activation at the cost of larger spectral deviations.
These results highlight the need to evaluate Earth-observation super-resolution beyond visual quality alone. For downstream geospatial applications, the relevant question is not only whether SR introduces additional spatial structure, but whether that structure remains consistent with the original spectral measurements and improves the derived thematic product.
The complete workflow, data, and generated maps are openly available.
Read the publication | Code and data
Explore the interactive results
The super-resolution outputs, spectral-index maps, and model comparisons for both case studies can be explored interactively below.













