适配器(计算)
计算机科学
人工智能
计算机视觉
图像(数学)
弹丸
计算机硬件
有机化学
化学
作者
Gang Du,Hanzi Wang,Xintao Xu,Yan Yan,Xuelong Li
标识
DOI:10.1109/tcsvt.2025.3602826
摘要
In recent years, few-shot image classification has achieved substantial progress. Although existing methods have achieved promising performance, the limited availability of training data often leads to the problem of model overfitting. Model overfitting affects generalization and restricts the effective transfer of knowledge to unseen classes. Moreover, existing methods maintain independence between the image and text modalities during the encoding process, lacking mutual collaboration. This limitation restricts their ability to fully exploit task-specific semantic relationships between visual concepts and textual descriptions. To address this challenge, we propose a text-driven cross-modal feature fusion adapter (TCFF-Adapter) for few-shot image classification. TCFF-Adapter introduces two core components: a cross-modal feature fusion module that constructs joint representations by aligning image and text semantics, and a text-driven adapter that optimizes fused features and dynamically adjusts feature weights in a meta-learning paradigm. By integrating multimodal knowledge with parameter-efficient tuning, our method achieves robust generalization to unseen data without requiring additional fine-tuning. Extensive experiments on eight benchmark datasets demonstrate that the proposed TCFF-Adapter significantly outperforms various state-of-the-art few-shot image classification methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI