计算机科学
人工智能
目标检测
Boosting(机器学习)
模式识别(心理学)
班级(哲学)
机器学习
强化学习
特征提取
对象(语法)
适应性
视觉对象识别的认知神经科学
特征(语言学)
分类器(UML)
一般化
启发式
约束(计算机辅助设计)
任务(项目管理)
上下文图像分类
排名(信息检索)
计算机视觉
缩小
钥匙(锁)
对象类检测
基础(拓扑)
任务分析
特征学习
透视图(图形)
最小边界框
火车
作者
Ruisong Zhang,Xin-Jian Wu,Changhu Wang,Cheng-Lin Liu
标识
DOI:10.1109/tmm.2026.3668516
摘要
Object detection, which aims to locate and recognize objects in images, is evolving toward reduced reliance on manual annotations and enhanced adaptability to open-world scenarios. This shift has led to open-vocabulary object detection (OVD), which enables zero-shot detection of objects from novel categories beyond the base categories. In this work, we identify three key challenges in detecting unseen class instances: 1) locating the instances of new classes; 2) distinguishing new class instances from the background; 3) recognizing new class instances. We propose a detection framework that leverages vision-language pre-trained (VLPT) models, such as CLIP, as the backbone to jointly address these three challenges. Specifically, we treat localization as a box-deformation decision process, where the agent interacts with the image to learn a universal deformation strategy, enhancing generalization for unseen class objects. We further reformulate the foreground-background classification as an objectness ranking task to improve objectness evaluation, utilizing a specially designed AP loss. Additionally, a feature magnitude minimization constraint is introduced for the adapter during fine-tuning, boosting recognition performance for both base and novel classes. Experiments on COCO and LVIS datasets demonstrate that our method outperforms previous approaches in open-vocabulary object detection.
科研通智能强力驱动
Strongly Powered by AbleSci AI