人工智能
计算机科学
卷积神经网络
变压器
闭塞
模式识别(心理学)
地点
计算机视觉
深度学习
视觉对象识别的认知神经科学
特征提取
医学
物理
哲学
量子力学
心脏病学
电压
语言学
作者
Jiseong Heo,Yooseung Wang,Jihun Park
标识
DOI:10.1016/j.patrec.2022.05.006
摘要
Object classification under partial occlusion has been challenging for deep convolutional neural networks due to their innate locality in extracting features. We propose an Occlusion-aware Spatial Attention Transformer (OSAT) architecture based on Vision Transformer (ViT), CutMix augmentation, and Occlusion Mask Predictor (OMP) to solve the occlusion problem. ViT mainly utilizes the self-attention mechanism, which enables the model to capture spatially distant information. In addition, for occluded image augmentation, we combine CutMix augmentation with ViT. OMP is used as a multi-task learning method and for spatial attention on non-occluded region. Our proposed OSAT achieves state-of-the-art performance on occluded vehicle classification datasets from PASCAL3D+ and MS-COCO. Moreover, additional experiments show that OMP outperforms previous approach in occluder localization both quantitatively and qualitatively. According to our ablation studies, ViT is effective at analyzing occluded objects, and our approach of CutMix augmentation and OMP led to further improvements.
科研通智能强力驱动
Strongly Powered by AbleSci AI