可解释性
雅卡索引
计算机科学
人工智能
机器学习
编码(社会科学)
纳克
模态(人机交互)
相似性(几何)
源代码
钥匙(锁)
编码(集合论)
自然语言处理
数据挖掘
语言模型
模式识别(心理学)
统计
计算机安全
数学
集合(抽象数据类型)
图像(数学)
程序设计语言
操作系统
作者
Keyang Xu,Mike Lam,Jingzhi Pang,Xin Gao,Charlotte Band,Piyush Mathur,Frank Papay,Ashish Khanna,Jacek B. Cywiński,Kamal Maheshwari,Pengtao Xie,Eric P. Xing
标识
DOI:10.48550/arxiv.1810.13348
摘要
This study presents a multimodal machine learning model to predict ICD-10 diagnostic codes. We developed separate machine learning models that can handle data from different modalities, including unstructured text, semi-structured text and structured tabular data. We further employed an ensemble method to integrate all modality-specific models to generate ICD-10 codes. Key evidence was also extracted to make our prediction more convincing and explainable. We used the Medical Information Mart for Intensive Care III (MIMIC -III) dataset to validate our approach. For ICD code prediction, our best-performing model (micro-F1 = 0.7633, micro-AUC = 0.9541) significantly outperforms other baseline models including TF-IDF (micro-F1 = 0.6721, micro-AUC = 0.7879) and Text-CNN model (micro-F1 = 0.6569, micro-AUC = 0.9235). For interpretability, our approach achieves a Jaccard Similarity Coefficient (JSC) of 0.1806 on text data and 0.3105 on tabular data, where well-trained physicians achieve 0.2780 and 0.5002 respectively.
科研通智能强力驱动
Strongly Powered by AbleSci AI