Prompt-Tuning-Guided Dual-Distribution Alignment for Unsupervised 2D–3D Cross-Modal Retrieval

计算机科学 判别式 特征学习 人工智能 特征(语言学) 特征提取 模式识别(心理学) 滤波器(信号处理) 情态动词 代表(政治) 合并(版本控制) 匹配(统计) 任务(项目管理) 机器学习 对偶(语法数字) 任务分析 联营 渲染(计算机图形) 特征匹配 语义学(计算机科学) 多任务学习
作者
Yaqian Zhou,Ruiqiang Guo,Dan Song,Jiayu Li,An-An Liu
出处
期刊:IEEE Transactions on Circuits and Systems for Video Technology [Institute of Electrical and Electronics Engineers]
卷期号:36 (7): 9197-9211
标识
DOI:10.1109/tcsvt.2026.3668410
摘要

2D-3D cross-modal retrieval (2D-3DCMR) aims at retrieving the most matching 3D models by leveraging 2D query images. However, the inherent modal discrepancy makes the 2D-3DCMR task still largely challenging. Besides, the scarcity of 3D labels in real-world applications severely hinders learning discriminative representations. To address these limitations, we present a prompt tuning guided dual distribution alignment (PTG-DDA) framework based on the CLIP model for the 2D-3DCMR task. Specifically, we design a learnable multi-view adaptive representation learning (MARL) module that adaptively integrates 3D features to merge complementary information and filter out redundant information across views, thereby improving the representation capability of 3D models. To mitigate the feature distribution shift between the 2D and 3D data, we design an attention-guided heterogeneous feature alignment (AHFA) module to guide the 2D and 3D inputs attend to feature banks by adopting the attention mechanism, thereby achieving heterogeneous feature alignment. Furthermore, to learn discriminative 3D features, we employ a multi-modal semantic prompt synergy (MSPS) module, which integrates class-related representations into learnable prompts to progressively learn the cross-modal synergy via a prompt synergy adapter, thereby achieving semantic feature alignment. Comprehensive experimental results on popular 2D-3DCMR benchmarks, i.e., MI3DOR and MI3DOR-2, demonstrate the superiority and effectiveness of PTG-DDA.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
wang发布了新的文献求助10
刚刚
刚刚
syangZ发布了新的文献求助10
刚刚
Diana发布了新的文献求助10
刚刚
李热热完成签到,获得积分10
1秒前
ijn发布了新的文献求助30
1秒前
海因伯顿完成签到,获得积分10
2秒前
白竹完成签到,获得积分10
2秒前
小坤完成签到,获得积分10
2秒前
2秒前
yyyy发布了新的文献求助10
2秒前
阳光金针菇完成签到 ,获得积分10
2秒前
LXD完成签到,获得积分10
2秒前
今后的应助被动听妙芙采纳,获得10
2秒前
Momo01发布了新的文献求助10
3秒前
3秒前
3秒前
wrong发布了新的文献求助10
3秒前
3秒前
1223完成签到,获得积分10
3秒前
3秒前
综述白完成签到,获得积分10
4秒前
4秒前
5秒前
6秒前
顾矜的应助被超级灰狼采纳,获得10
6秒前
爆米花的应助被hhhhhh采纳,获得10
6秒前
Junlin完成签到,获得积分10
6秒前
小蘑菇的应助被renyi采纳,获得30
6秒前
cuican完成签到 ,获得积分10
6秒前
Fuseme完成签到,获得积分10
7秒前
htt发布了新的文献求助10
7秒前
7秒前
Cyrus的应助被yang采纳,获得10
7秒前
7秒前
花灯王子发布了新的文献求助10
7秒前
qd完成签到,获得积分10
7秒前
Xio发布了新的文献求助10
8秒前
8秒前
所所的应助被Diana采纳,获得10
8秒前
高分求助中
(应助此贴封号)通过应助OA文献获取积分 10000
CODESSA Version 2.13 for Windows 2000
Rosenblum, Global Change Biology 800
Organizational Behavior 510
A Silent Apostrophe:The Fayum Portraits 350
Sing with Understanding: Introduction to Theology in Christian Congregational Song, 3rd ed 330
Protection enhancement strategies of potential outbreaks during Hajj 300
热门求助领域 (近24小时)
化学 材料科学 医学 生物 计算机科学 工程类 纳米技术 有机化学 化学工程 内科学 物理 生物化学 复合材料 催化作用 细胞生物学 人工智能 心理学 无机化学 基因 遗传学
热门帖子
关注 科研通微信公众号,转发送积分 7843945
求助须知:如何正确求助?哪些是违规求助? 9364642
关于积分的说明 20641063
捐赠科研通 7439650
什么是DOI,文献DOI怎么找? 3341003
关于科研通互助平台的介绍 2484952
邀请新用户注册赠送积分活动 2363099