清晨好,您是今天最早来到科研通的研友!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您科研之路漫漫前行!

Using a Longformer Large Language Model for Segmenting Unstructured Cancer Pathology Reports

病理 市场细分 计算机科学 自然语言处理 医学 人工智能 业务 营销
作者
Didier Fung,Gregory Arbour,Karenina Nurmelita Malik,Kaitlin Muzio,Raymond T. Ng
出处
期刊:JCO clinical cancer informatics [Lippincott Williams & Wilkins]
卷期号:9 (9): e2400143-e2400143
标识
DOI:10.1200/cci-24-00143
摘要

PURPOSE Many Natural Language Processing (NLP) methods achieve greater performance when the input text is preprocessed to remove extraneous or unnecessary text. A technique known as text segmentation can facilitate this step by isolating key sections from a document. Give that transformer models—such as Bidirectional Encoder Representations from Transformers (BERT)—have demonstrated state-of-the-art performance on many NLP tasks, it is desirable to leverage such models for segmentation. However, transformer models are typically limited to only 512 input tokens and are not well suited for lengthy documents such as cancer pathology reports. The Longformer is a modified transformer model designed to intake longer documents while retaining the positive characteristics of standard transformers. This study presents a Longformer model fine-tuned for cancer pathology report segmentation. METHODS We fine-tuned a Longformer Question-Answer (QA) model on 504 manually annotated pathology reports to isolate sections such as diagnosis, addenda, and clinical history. We compared baseline methods including regular expressions (regex) and BERT QA. However, those methods may fail to correctly identify section boundaries. Model performance was evaluated using sequence recall, precision, and F1 score. RESULTS Final test results were obtained on a hold-out test set of 304 cancer pathology reports. We report sequence F1 scores for the following sections: diagnosis (0.77), addenda (0.48), clinical history (0.89), and overall (0.68). CONCLUSION We present a fine-tuned Longformer model to isolate key sections from cancer pathology reports for downstream analyses. Our model performs segmentation with greater accuracy.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
Re完成签到 ,获得积分10
5秒前
10秒前
Pami发布了新的文献求助10
14秒前
冷酷冬卉完成签到,获得积分10
15秒前
传奇3应助Pami采纳,获得10
15秒前
37秒前
39秒前
42秒前
43秒前
彩色雪糕发布了新的文献求助10
44秒前
46秒前
CipherSage应助许多知识采纳,获得10
47秒前
Pami发布了新的文献求助10
47秒前
51秒前
54秒前
彩色雪糕完成签到,获得积分10
54秒前
CipherSage应助彩色雪糕采纳,获得10
1分钟前
自然的妙梦完成签到,获得积分10
1分钟前
缓慢的映天完成签到,获得积分10
1分钟前
Youy完成签到 ,获得积分10
1分钟前
1分钟前
难过的耳机完成签到,获得积分10
2分钟前
洁净马里奥完成签到 ,获得积分10
2分钟前
tlh完成签到 ,获得积分10
2分钟前
传奇3应助科研通管家采纳,获得30
2分钟前
乐乐应助科研通管家采纳,获得10
2分钟前
星辰大海应助科研通管家采纳,获得10
2分钟前
3分钟前
aa发布了新的文献求助10
3分钟前
3分钟前
3分钟前
贤惠的觅夏完成签到,获得积分10
3分钟前
Ava应助赵铁皮采纳,获得10
3分钟前
MchemG完成签到,获得积分0
3分钟前
aa完成签到,获得积分10
3分钟前
3分钟前
阿洁发布了新的文献求助10
3分钟前
3分钟前
黄锐完成签到 ,获得积分10
3分钟前
斯文败类应助阿洁采纳,获得10
3分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
The anomeric effect 1000
Principles of town planning: translating concepts to applications 1000
1 Peter and Christ's Descent to the Dead in Its Early Christian Reception 700
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7732482
求助须知:如何正确求助?哪些是违规求助? 9283233
关于积分的说明 20156465
捐赠科研通 7309908
什么是DOI,文献DOI怎么找? 3304109
关于科研通互助平台的介绍 2456922
邀请新用户注册赠送积分活动 2313258