Softmax函数
计算机科学
命名实体识别
法律文件
自然语言处理
子序列
情报检索
特征(语言学)
人工智能
基线(sea)
法律案件
语言学
数学
卷积神经网络
工程类
任务(项目管理)
哲学
系统工程
法学
数学分析
有界函数
政治学
作者
Xinrui Zhang,Xudong Luo,Jiaye Wu
标识
DOI:10.1109/ijcnn54540.2023.10191275
摘要
In legal practice, judicial professionals often need to extract useful information from numerous legal documents. The key to legal information extraction is the Named Entity Recognition (NER) of legal documents. To address three critical issues in NER of legal documents, this paper proposes a RoBERTa-GlobalPointer-based method for NER of legal documents. Specifically, we first use RoBERTa (a variant of the pre-trained language model BERT) to extract char-level feature representations of a legal document, and use the Skip-Gram method to extract its word-level feature representations, and fuse them to better capture the contextual information of entities in the document. Then, according to the concatenated result, we use the GlobalPointer method to calculate the score of each subsequence of the document, to which it is an entity of a certain type. Finally, we employ the balanced softmax function to determine whether or not a subsequence of the document is an entity of a certain type according to its score calculated by GlobalPointer. Our evaluation experiments on the Chinese judicial domain dataset show that the proposed method outperforms the state-of-the-art baseline methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI