自然语言处理
计算机科学
人工智能
任务(项目管理)
学期
判决
语义相似性
语言学
等价(形式语言)
相似性(几何)
特征(语言学)
词汇
选择(遗传算法)
语义学(计算机科学)
语义特征
深层语言处理
程序设计语言
管理
经济
哲学
图像(数学)
作者
Chunlin Wang,Irene Castellón Masalles,Elisabet Comelles
摘要
Abstract Semantic Textual Similarity (STS), which measures the equivalence of meanings between two textual segments, is an important and useful task in Natural Language Processing. In this article, we have analyzed the datasets provided by the Semantic Evaluation (SemEval) 2012–2014 campaigns for this task in order to find out appropriate linguistic features for each dataset, taking into account the influence that linguistic features at different levels (e.g. syntactic constituents and lexical semantics) might have on the sentence similarity. Results indicate that a linguistic feature may have a different effect on different corpus due to the great difference in sentence structure and vocabulary between datasets. Thus, we conclude that the selection of linguistic features according to the genre of the text might be a good strategy for obtaining better results in the STS task. This analysis could be a useful reference for measuring system building and linguistic feature tuning.
科研通智能强力驱动
Strongly Powered by AbleSci AI