Stage-Adaptive Spatial-Frequency Decomposition and Enhancement With Multimodal Conditional Routing for Remote Sensing Image Segmentation
Xuran Pan, Rui Zhang, Zibo Xu, Tingting Zhao
Tianjin University of Science and Technology
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Multimodal remote sensing image segmentation benefits from complementary cues provided by heterogeneous data sources, but accurate segmentation remains challenging due to difficulties in effectively fusing multimodal features and preserving fine-grained structures. Although spatial-frequency fusion methods have shown promise, existing approaches often rely on fixed or input-agnostic frequency partitioning and overlook the different objectives of encoding and decoding stages, limiting their adaptability to scene-dependent frequency distributions and stage-specific representation needs. To address these limitations, we propose a stage-adaptive spatial-frequency decomposition and enhancement network (SF-DENet) for multimodal remote sensing image segmentation. SF-DENet adopts a two-stream encoder–decoder architecture and introduces two stage-specific spatial-frequency fusion modules with distinct interaction patterns. During encoding, the encoder spatial-frequency fusion (EnFusion) module feeds the fused optical–auxiliary features into parallel spatial and frequency branches and employs adaptive frequency masks to modulate the amplitude spectrum, thereby enhancing cross-modal semantic alignment. During decoding, the decoder spatial-frequency fusion (DeFusion) module introduces explicit high- and low-frequency cues guided by the original inputs and performs frequency-aware interaction between decoder features and skip-connected encoder features, thereby enhancing structure-aware detail reconstruction. To overcome fixed or input-agnostic frequency partitioning, the frequency branch incorporates a conditional composition-based frequency decomposition (C $^{2}$ FD) module, which predicts routing weights from multimodal inputs and composes multiple soft-edge frequency-mask experts into content-adaptive masks for frequency representation modulation. Experiments on multiple multimodal remote sensing benchmarks demonstrate that SF-DENet achieves superior segmentation performance, particularly in scenes with complex textures, ambiguous boundaries, and modality inconsistency.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
物理Seismic Imaging and Inversion Techniques
Sparse and Compressive Sensing Techniques · Remote-Sensing Image Classification
参考文献 50
此处列出前 3 条