Bakar, Aqsa Abu and Khoshbakht, Amirreza and Aptoula, Erchan (2025) GeoSMAViT: geographical-context spectral-invariant multi-scale adaptive vision-language transformer. In: 15th Workshop on Hyperspectral Imaging and Signal Processing: Evolution in Remote Sensing (WHISPERS), Barcelona, Spain
Full text not available from this repository. (Request a copy)
Official URL: https://dx.doi.org/10.1109/WHISPERS69515.2025.11501561
Abstract
Cross-domain hyperspectral image classification faces challenges from deficiencies in geographical context modeling, spectral instabilities, and suboptimal multimodal fusion. We present GeoSMAViT, a geography-aware vision-language transformer with: (1) climate-aware spectral processing using region-specific band weighting and atmospheric correction, (2) multi-scale regional pattern recognition for different geographic land-use types, (3) single-stage cross-modal fusion with multi-head attention between vision and geographic text features, (4) progressive domain adaptation using source features for invariant representation, and comprehensive anti-overfitting measures including dynamic regularization and adaptive text dropout. On Houston 2013-to-2018 and PaviaU-to-PaviaC, GeoSMAViT achieved 98.37% and 96.29% OA, respectively, surpassing state-of-the-art baselines. Ablations confirmed that geographic awareness provided the largest contribution (+38.33 %) to cross-domain generalization.
| Item Type: | Papers in Conference Proceedings |
|---|---|
| Uncontrolled Keywords: | Cross-Domain Adaptation; Geographical Context Integration; Hyperspectral Image Classification; Spectral Invariance; Vision-Language Transformers |
| Divisions: | Faculty of Engineering and Natural Sciences |
| Depositing User: | Erchan Aptoula |
| Date Deposited: | 03 Jul 2026 12:12 |
| Last Modified: | 03 Jul 2026 12:12 |
| URI: | https://research.sabanciuniv.edu/id/eprint/54169 |

