人工智能
计算机科学
模式识别(心理学)
像素
分割
特征提取
卷积神经网络
图像分割
计算机视觉
特征(语言学)
树(集合论)
代表(政治)
集合(抽象数据类型)
尺度不变特征变换
数学
政治
数学分析
哲学
语言学
政治学
程序设计语言
法学
作者
Clément Farabet,Camille Couprie,Laurent Najman,Yann LeCun
标识
DOI:10.1109/tpami.2012.231
摘要
Scene labeling consists of labeling each pixel in an image with the category of the object it belongs to. We propose a method that uses a multiscale convolutional network trained from raw pixels to extract dense feature vectors that encode regions of multiple sizes centered on each pixel. The method alleviates the need for engineered features, and produces a powerful representation that captures texture, shape, and contextual information. We report results using multiple postprocessing methods to produce the final labeling. Among those, we propose a technique to automatically retrieve, from a pool of segmentation components, an optimal set of components that best explain the scene; these components are arbitrary, for example, they can be taken from a segmentation tree or from any family of oversegmentations. The system yields record accuracies on the SIFT Flow dataset (33 classes) and the Barcelona dataset (170 classes) and near-record accuracy on Stanford background dataset (eight classes), while being an order of magnitude faster than competing approaches, producing a $(320\times 240)$ image labeling in less than a second, including feature extraction.
科研通智能强力驱动
Strongly Powered by AbleSci AI