Extracting Diagnostic Data from Unstructured Bone Marrow Biopsy Reports of Myeloid Neoplasms Utilizing a Customized Natural Language Processing (NLP) Algorithm

髓系白血病 医学 活检 骨髓 髓样 金标准(测试) 算法 病理 白血病 放射科 计算机科学 内科学
作者
Isaac Kunz,Ananth Peddinti,Tina Nguyen,Morgan Ward,Alexandra Asay,Michael W. Deininger,Srinivas K. Tantravahi,Samir Courdy
出处
期刊:Blood [Elsevier BV]
卷期号:132 (Supplement 1): 2272-2272 被引量:3
标识
DOI:10.1182/blood-2018-99-119049
摘要

Abstract Introduction:Valuable research data is limited in its use when it is unstructured and not stored in discrete meaningful fields. Reports of the bone marrow (BM) aspirate and biopsies performed in patients with suspected or confirmed myeloid neoplasms typically include blood counts, peripheral blood (PB) and BM aspirate/touch preparation differential counts, morphological interpretation of aspirate and core biopsy and ancillary data such as karyotype, fluorescent in situ hybridization (FISH) and molecular mutations. Final BM reports are typically reported in a semi-structured document that are sufficient for a single patient review but inadequate for large scale queries to identify patients with a specific diagnosis or capture important diagnostic data. Manual extraction of these fields is expensive, time consuming and error prone. The aim of this study is to develop a customized algorithm for automated extraction of data from bone marrow biopsy reports and generate a framework that allows us to perform large-scale queries. Methods:We randomly identified 148 patients with a diagnosis of a myeloid neoplasm: chronic myeloid leukemia (n=45), chronic myelomonocytic leukemia (n=54) and acute myeloid leukemia (n=57). Seven patients included in this analysis were initially diagnosed as CMML and subsequently transformed to acute myeloid leukemia. Total number of reports evaluated was 524. Numerical and text diagnostic data were extracted manually from the entire cohort selected, which is considered a gold standard. A customized rule based algorithm was developed for each data attribute using Natural Language Processing (I2E Text Mining platform, Linguamatics Ltd, Cambridge, UK). Numerical data captured included differential counts from peripheral blood, bone marrow aspirate or touch preparation. Diagnostic data was captured as included diagnostic interpretation of peripheral blood smear and bone marrow aspirate. The algorithms for extracting the data were previously trained on a separate cohort. Precision and recall calculated for each data attribute utilizing R programing language and statistical computing environment. The calculation of precision can be defined as an index to measure the accuracy or closeness of a measured value to a known value (also known as positive predictive value). Recall can be defined as a measure of ability to capture all data points of interest (true positive rate or sensitivity). F-measure combines precision and recall as a harmonic mean. Results:Overall accuracy for the data captured was precision n = 0.9117 and recall n =0.7951. Precision and recall values for numerical and text data is reported in Table 1 and Figure 1. Conclusion:Extraction of relevant diagnostic data from unstructured bone marrow biopsy reports through automated approach is feasible and accurate. This method saves time and can be utilized for automated extraction of unstructured pathology reports from patients with different hematologic malignancies. Capturing data and storing in structured formats will allow researchers to perform large-scale queries. At the Huntsman Cancer Institute, this data is stored in easily accessible database and linked to other databases such as tissue banking. This approach will allow physicians and translational researchers to find samples with specific diagnosis or molecular mutation, for example identifying AML patients with mutated FLT3 gene. Data on extraction of karyotype, FISH and molecular mutations is being analyzed for accuracy and will be presented at the meeting. Future work involves identifying and improving accuracy and expanding the algorithms to extract additional fields in bone marrow biopsies and apply these algorithms to other hematologic malignancies. Disclosures Deininger: Blueprint: Consultancy; Pfizer: Consultancy, Membership on an entity's Board of Directors or advisory committees.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
sml的应助被太叔开山采纳,获得10
刚刚
bkagyin的应助被12138采纳,获得10
刚刚
2秒前
2秒前
wanci的应助被kela采纳,获得30
3秒前
xing_xing的应助被友好的季节采纳,获得20
3秒前
3秒前
陈业祝关注了科研通微信公众号
3秒前
4秒前
5秒前
bi8bo发布了新的文献求助10
6秒前
小豆包发布了新的文献求助10
7秒前
玉米玉米发布了新的文献求助10
7秒前
太叔开山完成签到,获得积分20
8秒前
鼎元发布了新的文献求助10
8秒前
呵呵哒发布了新的文献求助10
8秒前
lazyp完成签到 ,获得积分10
9秒前
Huihui完成签到,获得积分10
10秒前
11秒前
11秒前
Nature发布了新的文献求助10
11秒前
ding的应助被酒酽春浓采纳,获得10
12秒前
天天快乐的应助被朴实醉柳采纳,获得10
13秒前
13秒前
桐桐的应助被陈业祝采纳,获得10
13秒前
木木夕云发布了新的文献求助10
14秒前
15秒前
科研通AI2S的应助被小九采纳,获得10
16秒前
78888发布了新的文献求助10
17秒前
彭于晏完成签到,获得积分10
17秒前
热心不凡发布了新的文献求助10
17秒前
我是老大的应助被li采纳,获得10
18秒前
万能图书馆的应助被钱笑采纳,获得10
20秒前
20秒前
啦啦完成签到 ,获得积分10
21秒前
我是老大的应助被syb采纳,获得10
22秒前
SciGPT的应助被诚心的雁采纳,获得10
23秒前
24秒前
lqcolleen发布了新的文献求助10
24秒前
25秒前
高分求助中
(应助此贴封号)通过应助OA文献获取积分 10000
Rosenblum, Global Change Biology 800
The Dawn of Philology 520
Organizational Behavior 510
Production Logging: Theoretical and Interpretive Elements 400
A primer on partial least squares structural equation modeling (PLS-SEM) (4th ed.) 310
中国器官捐献和移植发展报告(2024) 300
热门求助领域 (近24小时)
化学 材料科学 医学 生物 计算机科学 工程类 纳米技术 有机化学 化学工程 内科学 物理 生物化学 复合材料 催化作用 细胞生物学 人工智能 心理学 无机化学 基因 遗传学
热门帖子
关注 科研通微信公众号,转发送积分 7822097
求助须知:如何正确求助?哪些是违规求助? 9348933
关于积分的说明 20550282
捐赠科研通 7414844
什么是DOI,文献DOI怎么找? 3333210
关于科研通互助平台的介绍 2479118
邀请新用户注册赠送积分活动 2353610