PaperScope
LIVE · 2026-10-07 05:40 UTC

How Many Independent Samples Does a Satellite Image Contain? Generalization Bounds for Spatially Dependent Data

Robin Young

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.08227 v1
Submitted
2026-10-06

Abstract

Machine learning classifiers for remote sensing imagery are typically evaluated as though every pixel were an independent sample. Spatial autocorrelation violates this assumption, since neighboring pixels carry redundant information which inflates sample sizes. How many independent samples does a satellite image actually contain? For an $n \times n$ image whose spatial correlation persists over a range of $r$ pixels, the effective sample size is $Θ(n^2/r^2)$, not $n^2$. We prove this as a finite-sample upper bound for classifiers on spatially correlated data, and show via a matching lower bound that the rate is tight, and no algorithm can do better. We extend the results to images with directional correlation and spatially varying correlation structure. Our result justifies spatial cross-validation since block holdout with separation proportional to the correlation range achieves optimal generalization guarantees, while random holdout can underestimate confidence interval widths by a factor proportional to $r$. We validate the theory on synthetic data and satellite image tiles from three sensors (Landsat 8, Sentinel-2, and Sentinel-1).

Journal: IEEE Transactions on Geoscience and Remote Sensing (2026) vol. 64

arXiv abs page · PDF · same-day batch