计算机科学
人工智能
模式识别(心理学)
卷积(计算机科学)
像素
比例(比率)
一致性(知识库)
代表(政治)
目标检测
特征(语言学)
特征学习
对象(语法)
卷积神经网络
计算机视觉
特征提取
人工神经网络
语言学
哲学
物理
量子力学
政治
政治学
法学
作者
Jingtao Xu,Yali Li,Shengjin Wang
标识
DOI:10.1109/cvidliccea56201.2022.9824076
摘要
It is challenging for convolution neural networks (CNN) to handle aerial images with extremely small objects. At each layer of the CNN, the limited receptive field leads to the contradiction between learning detailed features of the small objects and the global features of the large objects. Moreover, it is difficult to learn the feature representation for small objects occupying quite a few pixels. In this letter, we exploit the cross-scale consistency to enhance the feature representation of the objects at variant scales. We devise a double-stream training pipeline to capture the correspondence between objects at different scales and improve the adaptation of receptive field to scale variation. To explicitly formulate the model adaptability to object scales, we propose a novel scale consistency loss to learn the scale-invariant feature representation for object detection. To verify the effectiveness of our method, extensive experiments are conducted on the VisDrone-DET dataset which has large variance in object scale and quite small objects. Our method achieves state-of-the-art performance, outperforming existing methods by about 2.50%.
科研通智能强力驱动
Strongly Powered by AbleSci AI