Deep learning for small object detection in images

作者
Yang Liu
标识
DOI:10.32469/10355/79470
摘要

[ACCESS RESTRICTED TO THE UNIVERSITY OF MISSOURI AT REQUEST OF AUTHOR.] With the rapid development of deep learning in computer vision, especially deep convolutional neural networks (CNNs), significant advances have been made in recent years on object recognition and detection in images. Highly accurate detection results have been achieved for large objects, whereas detection accuracy on small objects remains to be low. This dissertation focuses on investigating deep learning methods for small object detection in images and proposing new methods with improved performance. First, we conducted a comprehensive review of existing deep learning methods for small object detections, in which we summarized and categorized major techniques and models, identified major challenges, and listed some future research directions. Existing techniques were categorized into using contextual information, combining multiple feature maps, creating sufficient positive examples, and balancing foreground and background examples. Methods developed in four related areas, generic object detection, face detection, object detection in aerial imagery, and segmentation, were summarized and compared. In addition, the performances of several leading deep learning methods for small object detection, including YOLOv3, Faster R-CNN, and SSD, were evaluated based on three large benchmark image datasets of small objects. Experimental results showed that Faster R-CNN performed the best, while YOLOv3 was a close second. Furthermore, a new deep learning method, called Retina-context Net, was proposed and outperformed state-of-the art one-stage deep learning models, including SSD, YOLOv3 and RetinaNet, on the COCO and SUN benchmark datasets. Secondly, we created a new dataset for bird detection, called Little Birds in Aerial Imagery (LBAI), from real-life aerial imagery. LBAI contains birds with sizes ranging from 10 by 10 pixels to 40 by 40 pixels. We adapted and applied several state-of-the-art deep learning models to LBAI, including object detection models such as YOLOv2, SSH, and Tiny Face, and instance segmentation models such as U-Net and Mask R-CNN. Our empirical results illustrated the strength and weakness of these methods, showing that SSH performed the best for easy cases, whereas Tiny Face performed the best for hard cases with cluttered backgrounds. Among small instance segmentation methods, U-Net achieved slightly better performance than Mask R-CNN. Thirdly, we proposed a new graph neural network-based object detection algorithm, called GODM, to take the spatial information of candidate objects into consideration in small object detection. Instead of detecting small objects independently as the existing deep learning methods do, GODM treats the candidate bounding boxes generated by existing object detectors as nodes and creates edges based on the spatial or semantic relationship between the candidate bounding boxes. GODM contains four major components: node feature generation, graph generation, node class labelling, and graph convolutional neural network model. Several graph generation methods were proposed. Experimental results on the LBDA dataset show that GODM outperformed existing state-of-the-art object detector Faster R-CNN significantly, up to 12% better in accuracy. Finally, we proposed a new computer vision-based grass analysis using machine learning. To deal with the variation of lighting condition, a two-stage segmentation strategy is proposed for grass coverage computation based on a blackboard background. On a real world dataset we collected from natural environments, the proposed method was robust to varying environments, lighting, and colors. For grass detection and coverage computation, the error rate was just 3%.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
1秒前
1秒前
mrx96完成签到 ,获得积分10
1秒前
2秒前
DOC_XIONG应助忧郁醉薇采纳,获得10
2秒前
jojo完成签到 ,获得积分10
2秒前
2秒前
3秒前
淡然的子默完成签到,获得积分10
3秒前
所所应助Man采纳,获得10
3秒前
4秒前
香蕉觅云应助隐形静芙采纳,获得10
4秒前
jiangjiaodu2024完成签到,获得积分10
4秒前
大个应助把妹王采纳,获得10
4秒前
faye发布了新的文献求助10
5秒前
6秒前
6秒前
标致梦露完成签到,获得积分10
7秒前
薇子完成签到,获得积分10
7秒前
那块发布了新的文献求助10
7秒前
洛廖琉发布了新的文献求助30
8秒前
乡非农卡完成签到 ,获得积分10
8秒前
斯文败类应助彩色炎彬采纳,获得10
8秒前
笨笨发布了新的文献求助10
8秒前
abc发布了新的文献求助10
9秒前
太阳alright发布了新的文献求助10
9秒前
66发布了新的文献求助10
9秒前
留胡子的如萱完成签到,获得积分10
10秒前
10秒前
12秒前
anny.white完成签到,获得积分10
12秒前
12秒前
su发布了新的文献求助10
12秒前
天天快乐应助太渊采纳,获得10
13秒前
hyx发布了新的文献求助10
14秒前
淡然的咖啡豆完成签到 ,获得积分10
14秒前
15秒前
anny.white发布了新的文献求助10
15秒前
KuangHS发布了新的文献求助30
15秒前
16秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Principles of town planning: translating concepts to applications 1000
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
核安全综合知识2024版 500
Photothermal Science and Techniques 500
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7718790
求助须知:如何正确求助?哪些是违规求助? 9272670
关于积分的说明 20093154
捐赠科研通 7294620
什么是DOI,文献DOI怎么找? 3299547
关于科研通互助平台的介绍 2453387
邀请新用户注册赠送积分活动 2306840