子序列
特征(语言学)
代表(政治)
拼音
词(群论)
自然语言处理
文字嵌入
人工智能
嵌入
计算机科学
相似性(几何)
语义学(计算机科学)
相关性(法律)
模式识别(心理学)
数学
汉字
语言学
图像(数学)
程序设计语言
几何学
法学
哲学
数学分析
政治
有界函数
政治学
作者
Yun Zhang,Yongguo Liu,Jiajing Zhu,Xindong Wu
标识
DOI:10.1109/taslp.2021.3073868
摘要
Chinese word embedding models capture Chinese semantics based on the character feature of Chinese words and the internal features of Chinese characters such as radical, component, stroke, structure and pinyin. However, some features are overlapping and most methods do not consider their relevance. Meanwhile, they express words as point vectors that cannot better capture different aspect semantics of Chinese words. In this paper, we propose a Feature Subsequence based Probability Representation Model (FSPRM) for learning Chinese word embeddings, in which we first integrate the morphological and phonetic features (stroke, structure and pinyin) of Chinese characters and learn their relevance by designing a feature subsequence to capture relatively comprehensive semantics of Chinese words, then feature probability distribution is proposed for capturing different aspect meanings of Chinese words based on the three internal features and probability representation by estimating its mean as the sum of feature subsequences. Chinese words with similar features may have similar semantics, then we map Chinese words to feature probability distributions and design a similarity-based objective for predicting the contextual words of the target word to learn their semantics. Extensive experiments on word analogy, word similarity, text classification and named entity recognition tasks demonstrate that the proposed method outperforms most state-of-the-art approaches.
科研通智能强力驱动
Strongly Powered by AbleSci AI