计算机科学
判别式
人工智能
分割
Boosting(机器学习)
模式识别(心理学)
适配器(计算)
过度拟合
卷积神经网络
辍学(神经网络)
棱锥(几何)
图像分割
阿达布思
计算机视觉
语音识别
特征提取
特征(语言学)
机器学习
一般化
稳健性(进化)
语义鸿沟
语义学(计算机科学)
特征学习
人工神经网络
深度学习
特征选择
尺度空间分割
作者
Changwei Wang,Wenhao Xu,Rongtao Xu,Zherui Zhang,Shibiao Xu,Jiguang Zhang,Xiaoqiang Teng,Weiliang Meng,Xiaopeng Zhang
标识
DOI:10.1109/tmm.2026.3654453
摘要
Open-vocabulary semantic segmentation is a challenging multimedia task that requires segmentation and recognition of unseen word classes during the testing phase. Recent works bridge the gap between closed and open-vocabulary recognition by introducing large-scale visual language models such as CLIP with cross-modal alignment capabilities. To preserve multimodal alignment capabilities, it is common to freeze the parameters of the CLIP and then add additional learnable components such as adapters to expand to downstream tasks. However, for the open-vocabulary semantic segmentation task, the plain adapter suffers from overfitting the closed-vocabulary classes and impairs performance on the open-vocabulary unseen classes. In addition, since CLIP is trained to perform image-level alignment can cause the network to over-focus on partially discriminative regions, resulting in incomplete segmentation masks. To alleviate the above problems, we introduce adaptive dropout adapters to release the Adaptive In Adapter (i.e. AIA) from the following two aspects: i) A Generalization Feature Selection Adapter (GFSA) is proposed to improve the generalization of network over unseen classes. ii) A Discriminative Region Mask Adapter (DRMA) is proposed for retrofitting CLIP backbone, has provided region free biased features for segmentation mask generation. Meanwhile, our proposed AIA achieves the current state-of-the-art performance on several open-vocabulary semantic segmentation benchmarks. Code is available at https://github.com/clearxu/AIA.
科研通智能强力驱动
Strongly Powered by AbleSci AI