修补
计算机科学
判别式
人工智能
分割
代表(政治)
财产(哲学)
城市景观
计算机视觉
对象(语法)
图像(数学)
模式识别(心理学)
艺术
哲学
视觉艺术
认识论
政治
政治学
法学
作者
Lingzhi Zhang,Tarmily Wen,Jie Min,Jiancong Wang,David K. Han,Jianbo Shi
标识
DOI:10.1007/978-3-030-58601-0_34
摘要
We study the problem of common sense placement of visual objects in an image. This involves multiple aspects of visual recognition: the instance segmentation of the scene, 3D layout, and common knowledge of how objects are placed and where objects are moving in the 3D scene. This seemingly simple task is difficult for current learning-based approaches because of the lack of labeled training pair of foreground objects paired with cleaned background scenes. We propose a self-learning framework that automatically generates the necessary training data without any manual labeling by detecting, cutting, and inpainting objects from an image. We propose a PlaceNet that predicts a diverse distribution of common sense locations when given a foreground object and a background scene. We show one practical use of our object placement network for augmenting training datasets by recomposition of object-scene with a key property of contextual relationship preservation. We demonstrate improvement of object detection and instance segmentation performance on both Cityscape [4] and KITTI [9] datasets. We also show that the learned representation of our PlaceNet displays strong discriminative power in image retrieval and classification.
科研通智能强力驱动
Strongly Powered by AbleSci AI