杠杆(统计)
计算机科学
稳健性(进化)
人工智能
特征(语言学)
融合
深度学习
特征提取
语义学(计算机科学)
传感器融合
机器学习
计算机视觉
滤波器(信号处理)
基线(sea)
模式识别(心理学)
语义特征
保险丝(电气)
卷积神经网络
数据挖掘
统计学习
马氏距离
作者
Chongchong Zan,Сонглин Ду,Takeshi Ikenaga
标识
DOI:10.23919/mva65244.2025.11175106
摘要
Pavement defect detection has drawn significant attention in computer vision research due to its indispensable role in infrastructure safety. Although deep learning has achieved huge success in daytime scenarios with abundant labeled data, its performance severely degrades under nighttime conditions due to insufficient defect samples and complex low-light noise. While few-shot learning alleviates data scarcity problem, existing approaches rely solely on visual features, ignoring the potential of cross-modal semantics to disambiguate defects from nighttime noise. Inspired by the success of vision-language models (VLMs) like Contrastive Language-Image Pre-Training (CLIP), this paper attempts to leverage text semantics to enhance robustness and proposes: (1) a Dual-Modal Hybrid Feature Extraction module which fully utilize ResNet local features, CLIP’s noise-robust visual features and text features from class-specific prompts, and (2) a Semantic-Guided Proposal Filter module that filter RPN proposals with low CLIP similarity to target-class text. Experiment results demonstrate that the proposed method outperforms baseline method, achieving a mean Average Precision (mAP) of 45.1% which surpasses vision-only baseline by 29.9% under 10-shot setting.
科研通智能强力驱动
Strongly Powered by AbleSci AI