GeoSMAViT: geographical-context spectral-invariant multi-scale adaptive vision-language transformer

Bakar, Aqsa Abu and Khoshbakht, Amirreza and Aptoula, Erchan (2025) GeoSMAViT: geographical-context spectral-invariant multi-scale adaptive vision-language transformer. In: 15th Workshop on Hyperspectral Imaging and Signal Processing: Evolution in Remote Sensing (WHISPERS), Barcelona, Spain

Full text not available from this repository. (Request a copy)

Abstract

Cross-domain hyperspectral image classification faces challenges from deficiencies in geographical context modeling, spectral instabilities, and suboptimal multimodal fusion. We present GeoSMAViT, a geography-aware vision-language transformer with: (1) climate-aware spectral processing using region-specific band weighting and atmospheric correction, (2) multi-scale regional pattern recognition for different geographic land-use types, (3) single-stage cross-modal fusion with multi-head attention between vision and geographic text features, (4) progressive domain adaptation using source features for invariant representation, and comprehensive anti-overfitting measures including dynamic regularization and adaptive text dropout. On Houston 2013-to-2018 and PaviaU-to-PaviaC, GeoSMAViT achieved 98.37% and 96.29% OA, respectively, surpassing state-of-the-art baselines. Ablations confirmed that geographic awareness provided the largest contribution (+38.33 %) to cross-domain generalization.
Item Type: Papers in Conference Proceedings
Uncontrolled Keywords: Cross-Domain Adaptation; Geographical Context Integration; Hyperspectral Image Classification; Spectral Invariance; Vision-Language Transformers
Divisions: Faculty of Engineering and Natural Sciences
Depositing User: Erchan Aptoula
Date Deposited: 03 Jul 2026 12:12
Last Modified: 03 Jul 2026 12:12
URI: https://research.sabanciuniv.edu/id/eprint/54169

Actions (login required)

View Item
View Item