计算机科学
构造(python库)
自然语言处理
人工智能
过程(计算)
多模态
绘画
语义学(计算机科学)
命名实体识别
融合
情报检索
图像融合
知识抽取
特征提取
可视化
答疑
人机交互
自然语言
语义记忆
接头(建筑物)
班级(哲学)
信息抽取
知识库
图像(数学)
作者
Jing Wan,Siyun Chen,Qingyang Zeng,Rumei Wang
出处
期刊:
[Springer Science+Business Media]
日期:2026-04-18
卷期号:14 (1)
标识
DOI:10.1038/s40494-026-02528-1
摘要
Chinese painting is an important form of cultural heritage. Its digital resources, including images and descriptive texts, contain rich yet unstructured information. Multimodal named entity recognition (MNER) is crucial for knowledge extraction but faces challenges including the lack of domain-specific datasets, limited integration of external knowledge, and semantic gaps between modalities. To address these issues, we construct CP-MNER, an MNER dataset for Chinese painting, and establish standardized baselines. We further propose MFKA, a multi-path fusion framework with knowledge augmentation. MFKA generates text-aware visual representations through cross-modal attention and incorporates external knowledge through a two-stage process using multimodal large language models. A multi-path complementary fusion module then integrates textual, visual, and knowledge-enhanced representations for multimodal semantic alignment. Experimental results demonstrate that MFKA achieves state-of-the-art performance on CP-MNER.
科研通智能强力驱动
Strongly Powered by AbleSci AI