计算机科学
判别式
特征学习
人工智能
特征(语言学)
特征提取
模式识别(心理学)
滤波器(信号处理)
情态动词
代表(政治)
合并(版本控制)
匹配(统计)
任务(项目管理)
机器学习
对偶(语法数字)
任务分析
联营
渲染(计算机图形)
特征匹配
语义学(计算机科学)
多任务学习
作者
Yaqian Zhou,Ruiqiang Guo,Dan Song,Jiayu Li,An-An Liu
标识
DOI:10.1109/tcsvt.2026.3668410
摘要
2D-3D cross-modal retrieval (2D-3DCMR) aims at retrieving the most matching 3D models by leveraging 2D query images. However, the inherent modal discrepancy makes the 2D-3DCMR task still largely challenging. Besides, the scarcity of 3D labels in real-world applications severely hinders learning discriminative representations. To address these limitations, we present a prompt tuning guided dual distribution alignment (PTG-DDA) framework based on the CLIP model for the 2D-3DCMR task. Specifically, we design a learnable multi-view adaptive representation learning (MARL) module that adaptively integrates 3D features to merge complementary information and filter out redundant information across views, thereby improving the representation capability of 3D models. To mitigate the feature distribution shift between the 2D and 3D data, we design an attention-guided heterogeneous feature alignment (AHFA) module to guide the 2D and 3D inputs attend to feature banks by adopting the attention mechanism, thereby achieving heterogeneous feature alignment. Furthermore, to learn discriminative 3D features, we employ a multi-modal semantic prompt synergy (MSPS) module, which integrates class-related representations into learnable prompts to progressively learn the cross-modal synergy via a prompt synergy adapter, thereby achieving semantic feature alignment. Comprehensive experimental results on popular 2D-3DCMR benchmarks, i.e., MI3DOR and MI3DOR-2, demonstrate the superiority and effectiveness of PTG-DDA.
科研通智能强力驱动
Strongly Powered by AbleSci AI