人工智能
计算机视觉
计算机科学
模式识别(心理学)
匹配(统计)
图像分割
图像(数学)
图像处理
特征提取
图像匹配
医学影像学
特征(语言学)
模式
迭代重建
上下文图像分类
图像配准
噪音(视频)
图像压缩
模板匹配
目标检测
图像去噪
作者
Xingyi He,Hao Yu,Sida Peng,Dongli Tan,Zehong Shen,Xiaowei Zhou,Hujun Bao
标识
DOI:10.1109/tpami.2026.3724652
摘要
Image matching, which aims to identify corresponding pixel locations between images, is crucial in a wide range of scientific disciplines, aiding in image registration, fusion, and analysis. However, when dealing with images captured under different imaging modalities that result in significant appearance changes, the performance of learning-based image matching algorithms often deteriorates due to the scarcity of annotated cross-modal training data. This limitation hinders applications in various fields that rely on multiple image modalities to obtain complementary information. To address this challenge, we propose a large-scale pre-training framework that utilizes synthetic cross-modal training signals, incorporating diverse data from various sources, to teach models to recognize and match fundamental structures across images. This capability is transferable to real-world, unseen cross-modality image matching tasks. Our key finding is that the matching model trained with our framework generalizes effectively across more than eight unseen cross-modality registration tasks using the same set of network weights, substantially outperforming existing generalizable methods and achieving competitive or superior performance compared to specialized models on several tasks.
科研通智能强力驱动
Strongly Powered by AbleSci AI