计算机科学
情态动词
空格(标点符号)
计算机视觉
人工智能
图像(数学)
计算机图形学(图像)
人机交互
操作系统
化学
高分子化学
作者
Yifan Jiao,Chenglong Cai,Bing‐Kun Bao
摘要
Unsupervised Domain Adaptation (UDA) aims to transfer models trained on a labeled source domain to an unlabeled target domain. Due to the excellent generalization ability of Vision Language Models (VLMs) such as CLIP in downstream tasks, most recent methods apply CLIP to UDA tasks through learning domain-specific text prompts for source and target domains separately. However, these methods fail to dynamically adjust image features based on the characteristics of their respective domains, thereby limiting their alignment with domain-specific text prompts in CLIP’s joint space, which is a key factor in improving classification performance in the target domain. To bridge this gap, we propose a Unified Text-Image Space Alignment with Cross-Modal Prompting (UTISA) framework for UDA. First, we introduce a Cross-Modal Prompt Learning (CMP) module to generate domain-specific image prompts and layer-specific image prompts for the visual branch to encode domain-specific knowledge globally and locally. Second, under the guidance of image prompts, we introduce a Domain-Aware Multi-Layer Feature Fusion (DMF) module to construct multi-layer domain features for each domain and enhance the image features with these multi-layer domain features, which enables the image features to better reflect the characteristics of their respective domains, thereby promoting their alignment with domain-specific text prompts. Moreover, we introduce a Perturbation-Driven Regularization (PDR) mechanism for the target domain to enhance the robustness and generalization of the model. The experiments demonstrate that UTISA achieves the best performance on three mainstream UDA benchmarks, including 87.9% on Office-Home, 90.9% on VisDA-2017, and 62.4% on DomainNet.
科研通智能强力驱动
Strongly Powered by AbleSci AI