计算机科学
人工智能
对象(语法)
语义学(计算机科学)
计算机视觉
自然语言处理
本体论
集合(抽象数据类型)
描述逻辑
语义数据模型
基于案例的推理
自动推理
知识表示与推理
可视化
作者
Peng Zhang,Xiaolong Wu,Yue Wang
标识
DOI:10.1109/iccsse67426.2025.11552579
摘要
Zero-Shot Object Navigation (ZSON) requires an embodied agent to locate a target in unseen environments without task-specific training. Existing vision-language approaches often rely on pixel-level cues or geometric partitions, limiting global reasoning and robustness in cluttered, dynamic, or occluded scenes. We propose GSR-Nav, a Graph-based Semantic Reasoning framework that incrementally constructs an online scene graph from RGB-D observations. Objects are represented as nodes with categories, attributes, and positions, while edges encode spatial and semantic relations. A relation-chain inference module ranks candidate goal locations by combining direct node-goal similarity with relational reasoning over unobserved targets, guiding a hierarchical planner that integrates global semantic context with local motion planning. We detail object detection, 3D localization, relation extraction, incremental updates, and inference weighting to ensure reproducibility. Experiments on HM3D and MP3D show that GSR-Nav surpasses state-of-the-art ZSON baselines in Success Rate (SR) and Success weighted by Path Length (SPL) [30]. Ablations and parameter analyses validate each module’s contribution, and tests in dynamic and occluded environments confirm the robustness of structured scene graph modeling.
科研通智能强力驱动
Strongly Powered by AbleSci AI