Spatial Multimodal Knowledge-Driven 3D Scene Graph Prediction With Vision-Language Model

计算机科学 场景图 人工智能 图形 空间关系 分层数据库模型 可视化 知识图 谓词(数理逻辑) 理论计算机科学 水准点(测量) 视觉语言 模式识别(心理学) 语义学(计算机科学) 图论 机器学习 语义映射 层级组织 利用 自然语言处理 邻接表 空间分析 空间语境意识
作者
H. S. Hou,Mingtao Feng,Zijie Wu,Yulan Guo,Yaonan Wang,Ajmal Mian
出处
期刊:IEEE Transactions on Circuits and Systems for Video Technology [Institute of Electrical and Electronics Engineers]
卷期号:36 (6): 7483-7498
标识
DOI:10.1109/tcsvt.2026.3664122
摘要

In-depth understanding of 3D environments not only involves locating and recognizing individual objects but also requires inferring the relationships and interactions among them. However, most existing methods heavily rely on scene-specific contents, which leads to poor performance due to the noisy, cluttered, and partial nature of real-world 3D scenes. In this work, we find that the inherently hierarchical structures of 3D environments, derived from support relationships, aid in the automatic association of semantic and spatial arrangements of objects and provide rich geometric and topological information independent of specific scenarios. To this end, we propose a 3D scene graph generation model that leverages the hierarchical structures of 3D environments as spatial multimodal knowledge to enhance 3D scene graph generation. Specifically, we first devise a cross-modal tuning approach, where a visually-prompted vision language model is learned to infer the support relationships between objects in a low-resource way. Subsequently, we build a hierarchical visual graph and hierarchical symbolic knowledge graph using the fine-tuned vision language model to extract contextualized visual contents and relevant textual facts, respectively. Finally, we progressively accumulate 3D spatial multimodal knowledge about the hierarchical structures by correlating contextualized visual contents and textual facts using a novel graph reasoning network. In addition, to better evaluate the performance of 3D scene graph generation models, we propose a new benchmark 3DSSG-M by reorganizing the widely-used 3D scene graph generation dataset 3DSSG. This reorganization balances the predicate distribution of 3DSSG and reduces the influence of frequency bias. Extensive results and ablations attest to the effectiveness of the hierarchical structures in 3D environments and demonstrate the superiority of our proposed method over current state-of-the-art competitors.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
刚刚
123发布了新的文献求助10
1秒前
cc发布了新的文献求助10
2秒前
3秒前
顺心不愁发布了新的文献求助10
3秒前
4秒前
cijing发布了新的文献求助10
5秒前
实验顺顺顺完成签到,获得积分10
5秒前
6秒前
6秒前
orixero应助机灵的海雪采纳,获得10
7秒前
7秒前
干八两螺蛳粉完成签到,获得积分10
8秒前
jjj完成签到,获得积分10
8秒前
CodeCraft应助123采纳,获得10
9秒前
10秒前
拟尼妮发布了新的文献求助10
10秒前
Wlc完成签到,获得积分10
11秒前
hang完成签到,获得积分10
11秒前
11秒前
hongw1980发布了新的文献求助10
11秒前
bai发布了新的文献求助30
12秒前
万能图书馆应助autumn采纳,获得10
12秒前
小二郎应助赵乂采纳,获得10
12秒前
只只发布了新的文献求助20
13秒前
科研通AI6.4应助行楽采纳,获得10
13秒前
13秒前
852应助瘦瘦的若山采纳,获得10
13秒前
14秒前
14秒前
wanci应助Forever采纳,获得10
15秒前
jackmilton发布了新的文献求助10
15秒前
FashionBoy应助興崋采纳,获得10
15秒前
初景应助甘甘甘甘甘采纳,获得20
15秒前
以念发布了新的文献求助10
16秒前
17秒前
领导范儿应助温婉的惜文采纳,获得10
17秒前
香蕉觅云应助QQ采纳,获得10
17秒前
路很遥远发布了新的文献求助10
17秒前
18秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
A Case Study on Hotels as Noncongregate Emergency Living Accommodations for Returning Citizens 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7765000
求助须知:如何正确求助?哪些是违规求助? 9309358
关于积分的说明 20310654
捐赠科研通 7349841
什么是DOI,文献DOI怎么找? 3314708
关于科研通互助平台的介绍 2464103
邀请新用户注册赠送积分活动 2329140