计算机科学
自然语言处理
模棱两可
变压器
人工智能
粒度
知识库
图形
自然语言
文本分割
分割
理论计算机科学
操作系统
物理
电压
程序设计语言
量子力学
作者
Boer Lyu,Lu Chen,Su Zhu,Kai Yu
标识
DOI:10.1609/aaai.v35i15.17592
摘要
Chinese short text matching is a fundamental task in natural language processing. Existing approaches usually take Chinese characters or words as input tokens. They have two limitations: 1) Some Chinese words are polysemous, and semantic information is not fully utilized. 2) Some models suffer potential issues caused by word segmentation. Here we introduce HowNet as an external knowledge base and propose a Linguistic knowledge Enhanced graph Transformer (LET) to deal with word ambiguity. Additionally, we adopt the word lattice graph as input to maintain multi-granularity information. Our model is also complementary to pre-trained language models. Experimental results on two Chinese datasets show that our models outperform various typical text matching approaches. Ablation study also indicates that both semantic information and multi-granularity information are important for text matching modeling.
科研通智能强力驱动
Strongly Powered by AbleSci AI