计算机科学
自然语言处理
人工智能
代表(政治)
领域(数学)
机器翻译
背景(考古学)
情绪分析
嵌入
命名实体识别
任务(项目管理)
生物
经济
法学
纯数学
政治学
政治
管理
数学
古生物学
作者
Rajvardhan Patil,Sorio Boit,Venkat N. Gudivada,Jagadeesh Nandigam
出处
期刊:IEEE Access
[Institute of Electrical and Electronics Engineers]
日期:2023-01-01
卷期号:11: 36120-36146
被引量:129
标识
DOI:10.1109/access.2023.3266377
摘要
Natural Language Processing (NLP) is a research field where a language in consideration is processed to understand its syntactic, semantic, and sentimental aspects. The advancement in the NLP area has helped solve problems in the domains such as Neural Machine Translation, Name Entity Recognition, Sentiment Analysis, and Chatbots, to name a few. The topic of NLP broadly consists of two main parts: the representation of the input text (raw data) into numerical format (vectors or matrix) and the design of models for processing the numerical data. This paper focuses on the former part and surveys how the NLP field has evolved from rule-based, statistical to more context-sensitive learned representations. For each embedding type, we list their representation, issues they addressed, limitations, and applications. This survey covers the history of text representations from the 1970s and onwards, from regular expressions to the latest vector representations used to encode the raw text data. It demonstrates how the NLP field progressed from where it could comprehend just bits and pieces to all the significant aspects of the text over time.
科研通智能强力驱动
Strongly Powered by AbleSci AI