遇见数据集

Deformable VMamba and feature alignment based UAV multimodal images semantic segmentation method (<italic>invited</italic>)

收藏
中国科学数据2026-04-24 更新2026-04-25 收录
官方服务:

资源简介:

ObjectiveVisible-infrared (RGB-T) multi-modal semantic segmentation is essential for UAV-based urban monitoring, intelligent patrol, and smart city management, leveraging thermal infrared data to compensate for visible images under poor illumination. However, existing methods face critical limitations: CNNs suffer from restricted receptive fields and weak global modeling; Transformer-based attention incurs excessive computational cost unsuitable for UAV edge platforms; and standard Mamba architectures lose 2D spatial structure information due to fixed scanning paths in visual tasks. Moreover, UAV oblique imaging introduces weak alignment between modalities, leading to feature fusion errors and degraded accuracy, particularly for small objects. To address these challenges, we proposes UAV-Mamba, a multi-modal semantic segmentation method based on deformable VMamba and cross-modal feature alignment, designed to enhance segmentation accuracy, robustness, and efficiency in weakly aligned UAV RGB-T scenes.MethodsWe propose a deformable VMamba and feature alignment based UAV multimodal images semantic segmentation method (Fig.3). First, a deformable scanning visual Mamba (D-VMamba) backbone (Fig.4) generates dynamic offsets to shift feature sampling from fixed grids to informative regions and adaptively adjusts scanning order, producing structure-aware feature sequences that align with object morphology and improve fine-grained perception. Second, a cross-modal feature alignment (CMFA) module (Fig.6) uses infrared features as a reference to learn a dynamic offset field from visible to infrared, achieving pixel-level adaptive registration and mitigating fusion errors caused by misalignment. Third, a Cross-Mamba interaction and adaptive fusion (CMIAF) module (Fig.7) enables bidirectional information exchange via cross selective scanning and adaptively weights fused features through a conflict-aware mechanism, fully exploiting multi-modal complementarity. Extensive experiments on the Kust4K dataset, including quantitative comparisons (Tab.1) with state-of-the-art (SOTA) methods, qualitative visual analysis (Fig.8), and ablation studies (Tab.2), validate the effectiveness of each module.Results and DiscussionsUAV-Mamba achieves 73.4% mIoU on Kust4K, outperforming the baseline model Sigma-T by 1.9% and the current SOTA model CMX-B4 by 0.5%, while maintaining superior efficiency with only 53.7M parameters and 100.5 GFLOPs, which are 40.2% and 70.0% of CMX-B4, respectively. For challenging small-object categories (Motorcycle, Person, Traffic Facilities), it achieves gains of 5.5%, 2.5%, and 1.1% over Sigma-T, demonstrating enhanced capability for small, occluded targets. Ablation studies confirm each module’s contribution. D-VMamba alone improves mIoU by 1.0% serving as the primary performance driver. Visual results further show that UAV-Mamba produces complete and precise segmentation under low illumination, occlusion, and dense small objects, consistent with quantitative findings.ConclusionsThis paper presents UAV-Mamba, a deformable VMamba-based framework with cross-modal alignment that effectively addresses weak modal alignment, spatial structure loss in standard Mamba, and insufficient feature fusion in UAV RGB-T segmentation. Achieving SOTA performance on Kust4K with strong robustness in complex urban environments, it offers reliable support for practical UAV applications including low-altitude patrol and intelligent traffic perception. Future work will focus on lightweight optimization to enable real-time segmentation in dynamic UAV flight scenarios.

创建时间:
2026-04-24
二维码
社区交流群
二维码
科研交流群
商业服务