计算机科学
人工智能
情态动词
推论
图形
链接(几何体)
可扩展性
光学(聚焦)
机器学习
编码(内存)
传感器融合
模式
任务(项目管理)
多模态
航程(航空)
模态(人机交互)
数据挖掘
融合
编码
多模式学习
领域知识
知识图
信息融合
领域(数学分析)
作者
Xiaodi Xu,Lijie Li,Ye Wang,Tao Ren,Tian Qiao
标识
DOI:10.1145/3746027.3755661
摘要
Multimodal link prediction on multimodal knowledge graphs is an inference task aimed at finding missing triples, which seeks to improve prediction accuracy by leveraging a wide range of information. However, current multimodal knowledge graph link prediction methods are primarily designed in the spatial domain, necessitating ever-growing complexity in fusion strategies. In addition, most of them focus on only three modalities (text, image, and structure). Forward-looking approaches, however, should accommodate a broader array of modalities. Motivated by the operational simplification enabled by transforming features into the frequency (or time-frequency) domain, we propose a wavelet-transform-based multimodal link prediction method, WFF, which offers high modal extensibility and low fusion complexity. Specifically, for unimodal information, we designed a Unimodal Time-Frequency Knowledge Enhancement module, UTFKE, which extracts time-frequency features via discrete wavelet transform and enhances information quality through adaptive filtering. To address the challenge of multimodal fusion, we devised a Multimodal Time-Frequency Knowledge Fusion module MTFKF that supports high modal extensibility and enables effective, efficient integration. Extensive experiments on multiple well-known datasets demonstrate that WFF outperforms strong baselines and achieves state-of-the-art performance. In addition, WFF extends modality to audio and video, further validating the model's effectiveness. Our code is available at https://github.com/xxd12315/WFF.
科研通智能强力驱动
Strongly Powered by AbleSci AI