An OCR Post-Correction Approach Using Deep Learning for Processing Medical Reports

光学字符识别 计算机科学 人工智能 深度学习 编码器 自然语言处理 过程(计算) 非结构化数据 领域(数学分析) 字错误率 情报检索 大数据 数据挖掘 图像(数学) 数学分析 数学 操作系统
作者
S Karthikeyan,Alba García Seco de Herrera,Faiyaz Doctor,Asim Mirza
出处
期刊:IEEE Transactions on Circuits and Systems for Video Technology [Institute of Electrical and Electronics Engineers]
卷期号:32 (5): 2574-2581 被引量:32
标识
DOI:10.1109/tcsvt.2021.3087641
摘要

According to a recent Deloitte study, the COVID-19 pandemic continues to place a huge strain on the global health care sector. Covid-19 has also catalysed digital transformation across the sector for improving operational efficiencies. As a result, the amount of digitally stored patient data such as discharge letters, scan images, test results or free text entries by doctors has grown significantly. In 2020, 2314 exabytes of medical data was generated globally. This medical data does not conform to a generic structure and is mostly in the form of unstructured digitally generated or scanned paper documents stored as part of a patient’s medical reports. This unstructured data is digitised using Optical Character Recognition (OCR) process. A key challenge here is that the accuracy of the OCR process varies due to the inability of current OCR engines to correctly transcribe scanned or handwritten documents in which text may be skewed, obscured or illegible. This is compounded by the fact that processed text is comprised of specific medical terminologies that do not necessarily form part of general language lexicons. The proposed work uses a deep neural network based self-supervised pre-training technique: Robustly Optimized Bidirectional Encoder Representations from Transformers (RoBERTa) that can learn to predict hidden (masked) sections of text to fill in the gaps of non-transcribable parts of the documents being processed. Evaluating the proposed method on domain-specific datasets which include real medical documents, shows a significantly reduced word error rate demonstrating the effectiveness of the approach.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
daisyee完成签到,获得积分10
刚刚
芋头发布了新的文献求助10
刚刚
WD发布了新的文献求助10
1秒前
科研通AI6.2应助李Tt采纳,获得10
1秒前
wanci应助zzzzzzz采纳,获得10
2秒前
小飞发布了新的文献求助10
2秒前
Cassiel关注了科研通微信公众号
2秒前
好学天上发布了新的文献求助10
3秒前
谦让友绿完成签到,获得积分10
3秒前
小栾发布了新的文献求助10
3秒前
3秒前
3秒前
MozzieMiao给会撒娇的哑铃的求助进行了留言
4秒前
云珀千完成签到 ,获得积分10
4秒前
4秒前
二十三号完成签到,获得积分20
4秒前
SciGPT应助纳米采纳,获得10
4秒前
5秒前
万能图书馆应助CC采纳,获得10
5秒前
Meria完成签到,获得积分10
5秒前
尊嘟假嘟发布了新的文献求助30
6秒前
科研通AI6.3应助edge采纳,获得10
6秒前
lin发布了新的文献求助10
6秒前
Anovel完成签到,获得积分10
7秒前
研友_LMpo68发布了新的文献求助10
7秒前
初晴发布了新的文献求助10
7秒前
蓝天发布了新的文献求助10
9秒前
9秒前
所所应助豆瓣酱采纳,获得10
10秒前
深情安青应助海洋球采纳,获得10
10秒前
林飞飞发布了新的文献求助10
10秒前
李子青发布了新的文献求助10
10秒前
韩书琴发布了新的文献求助10
11秒前
woods完成签到,获得积分10
11秒前
Avalonx应助ale采纳,获得10
12秒前
研究生发布了新的文献求助10
12秒前
12秒前
二十三号发布了新的文献求助10
12秒前
赘婿应助Bokuto采纳,获得10
13秒前
维克特瑞完成签到,获得积分10
13秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Nondestructive Testing Handbook: Vol. 4, Thermal and Infrared Testing (IR), 4th ed 800
作者名:Kristopher P. Plain,悉尼大学的,目前只能查到其四篇论文,想找到其博士论文 590
Évora na Idade Média 555
Soil mites of the family Rhagidiidae (Actinedida: Eupodoidea). Morphology, Systematics, Ecology 520
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Stratospheric Ozone: A Textbook 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7359344
求助须知:如何正确求助?哪些是违规求助? 8969418
关于积分的说明 19062361
捐赠科研通 7006214
什么是DOI,文献DOI怎么找? 3222853
关于科研通互助平台的介绍 2386786
邀请新用户注册赠送积分活动 2203685