计算机科学
渲染(计算机图形)
人工智能
语义学(计算机科学)
概率逻辑
计算机视觉
代表(政治)
分割
语义映射
钥匙(锁)
多样性(控制论)
几何本原
机器人
可视化
稳健性(进化)
图像分割
桥接(联网)
运动规划
实体造型
光流
传感器融合
几何造型
中间语言
语义数据模型
机器人学
人机交互
作者
Yinan Deng,Yufeng Yue,Jianyu Dou,Jingyu Zhao,Jiahui Wang,Yujie Tang,Yi Yang,Mengyin Fu
标识
DOI:10.1109/tro.2025.3621333
摘要
Robotic systems demand accurate and comprehensive 3D environment perception, requiring simultaneous capture of photo-realistic appearance (optical), precise layout shape (geometric), and open-vocabulary scene understanding (semantic). Existing methods typically achieve only partial fulfillment of these requirements while exhibiting optical blurring, geometric irregularities, and semantic ambiguities. To address these challenges, we propose OmniMap. Overall, OmniMap represents the first online mapping framework that simultaneously captures optical, geometric, and semantic scene attributes while maintaining real-time performance and model compactness. At the architectural level, OmniMap employs a tightly coupled 3DGS-Voxel hybrid representation that combines fine-grained modeling with structural stability. At the implementation level, OmniMap identifies key challenges across different modalities and introduces several innovations: adaptive camera modeling for motion blur and exposure compensation, hybrid incremental representation with normal constraints, and probabilistic fusion for robust instance-level understanding. Extensive experiments show OmniMap's superior performance in rendering fidelity, geometric accuracy, and zero-shot semantic segmentation compared to state-of-the-art methods across diverse scenes. The framework's versatility is further evidenced through a variety of downstream applications, including multi-domain scene Q&A, interactive editing, perception-guided manipulation, and map-assisted navigation.
科研通智能强力驱动
Strongly Powered by AbleSci AI