自编码
计算机科学
人工智能
特征(语言学)
背景(考古学)
语义特征
阶段(地层学)
比例(比率)
钥匙(锁)
自然语言处理
计算机视觉
模式识别(心理学)
深度学习
地图学
地理
语言学
生物
古生物学
计算机安全
哲学
考古
作者
Sangmin Park,Jong-Eun Ha
出处
期刊:제어로봇시스템학회 논문지
[Institute of Control, Robotics and Systems]
日期:2023-12-13
卷期号:29 (12): 966-972
标识
DOI:10.5302/j.icros.2023.23.0143
摘要
Autonomous systems require a profound understanding of their surroundings, encompassing both semantic and 3D geometry. This study focuses on advancing 3D semantic scene completion approaches using a camera. Building upon the foundation laid by VoxFormer [1], which is recognized for its state-of-the-art performance in 3D semantic scene completion, our approach involves two distinct stages. In the initial stage, scene completion is done with depth images, while in the second stage, the final 3D scene completion is performed using masked autoencoder. To enhance the performance of VoxFormer, we introduced two key modifications. First, we modified the first stage using multi-scale feature maps. Second, we further modified the first stage using a masked autoencoder. Experimental results, based on the adapted VoxFormer model in both stages are presented. Our two proposed approaches exhibit notable improvements, particularly in the context of small objects. However, these enhancements warrant further investigation for optimization and refinement.
科研通智能强力驱动
Strongly Powered by AbleSci AI