计算机科学
人工智能
计算机视觉
模式识别(心理学)
特征提取
图像匹配
特征匹配
特征(语言学)
图像(数学)
图像处理
变压器
图像分割
像素
模板匹配
特征检测(计算机视觉)
匹配(统计)
合成孔径雷达
算法
数据压缩
特征跟踪
迭代重建
图像配准
雷达成像
作者
Chenzhong Gao,Yunhao Gao,Daping Weng,Jixuan Li,Wei Li,Ran Tao,Xiang-Gen Xia
标识
DOI:10.1109/tgrs.2026.3683096
摘要
Image matching is a fundamental process in intelligent image processing, especially for the wide application of multi-source remote sensing technologies, where invariant feature extraction is a core technology. Deep learning-based matching networks have emerged in recent years due to their powerful perception capabilities and parallel computing efficiency. However, most of them are designed for optical or a specific modality. Only recently general cross-modal matching has gained attention, but the prevailing approaches mainly rely on fine-tuning existing models or utilizing frozen weights from large pre-trained models, with little focus on dedicated models tailored for cross-modal scenarios. In this paper, we propose Modality-Invariant Feature Transformer (MIFTr), a lightweight image matching model dedicated to achieving universal cross-modal capability. The model innovatively introduces a rotational multi-branch backbone, and incorporates the latest MambaVision Mixer to construct a Symmetric Cross-Vision-Mamba (CrossVM) module as the core interactive feature encoder. With sufficient pretraining, the model handles general types of image modalities with processing time within 40ms per 5122sample on RTX4090. Through extensive comparative experiments on several complex cross-modal datasets, the proposed MIFTr demonstrates consistently superior performance on each modality, while maintaining a minimal model size and exceptional computational efficiency. The codes and experimental data will be made publicly available athttps://github.com/MrPingQi/MIFTr.
科研通智能强力驱动
Strongly Powered by AbleSci AI