Clinical named entity recognition and relation extraction using natural language processing of medical free text: A systematic review

关系抽取命名实体识别计算机科学信息抽取人工智能自然语言处理生物医学文本挖掘机器学习情报检索关系（数据库）条件随机场领域（数学）计算语言学任务（项目管理）文本挖掘数据挖掘经济管理数学纯数学

作者

David Fraile Navarro,Kiran Ijaz,Dana Rezazadegan,Hania Rahimi-Ardabili,Mark Dras,Enrico Coiera,Shlomo Berkovsky

出处

期刊：International Journal of Medical Informatics [Elsevier]
日期：2023-09-01 卷期号：177: 105122-105122 被引量：3

链接

nih.govdoi.org

标识

DOI：10.1016/j.ijmedinf.2023.105122

摘要

Natural Language Processing (NLP) applications have developed over the past years in various fields including its application to clinical free text for named entity recognition and relation extraction. However, there has been rapid developments the last few years that there's currently no overview of it. Moreover, it is unclear how these models and tools have been translated into clinical practice. We aim to synthesize and review these developments.We reviewed literature from 2010 to date, searching PubMed, Scopus, the Association of Computational Linguistics (ACL), and Association of Computer Machinery (ACM) libraries for studies of NLP systems performing general-purpose (i.e., not disease- or treatment-specific) information extraction and relation extraction tasks in unstructured clinical text (e.g., discharge summaries).We included in the review 94 studies with 30 studies published in the last three years. Machine learning methods were used in 68 studies, rule-based in 5 studies, and both in 22 studies. 63 studies focused on Named Entity Recognition, 13 on Relation Extraction and 18 performed both. The most frequently extracted entities were "problem", "test" and "treatment". 72 studies used public datasets and 22 studies used proprietary datasets alone. Only 14 studies defined clearly a clinical or information task to be addressed by the system and just three studies reported its use outside the experimental setting. Only 7 studies shared a pre-trained model and only 8 an available software tool.Machine learning-based methods have dominated the NLP field on information extraction tasks. More recently, Transformer-based language models are taking the lead and showing the strongest performance. However, these developments are mostly based on a few datasets and generic annotations, with very few real-world use cases. This may raise questions about the generalizability of findings, translation into practice and highlights the need for robust clinical evaluation.

求助该文献

最长约 10秒，即可获得该文献文件

Clinical named entity recognition and relation extraction using natural language processing of medical free text: A systematic review

今日热心研友