计算机科学
人工智能
机器学习
深度学习
水准点(测量)
仿形(计算机编程)
人工神经网络
深层神经网络
标记数据
编码(社会科学)
生物学数据
领域知识
图形
线性模型
数据挖掘
随机森林
支持向量机
基因表达谱
领域(数学分析)
领域(数学)
基因调控网络
数据建模
噪声数据
遗传程序设计
可解释性
基因预测
监督学习
作者
Matthew B.A. McDermott,Jennifer Wang,Wen-Ning Zhao,Steven D. Sheridan,Peter Szolovits,Isaac Kohane,Stephen J. Haggarty,Roy H. Perlis
标识
DOI:10.1109/tcbb.2019.2910061
摘要
Gene expression data can offer deep, physiological insights beyond the static coding of the genome alone. We believe that realizing this potential requires specialized, high-capacity machine learning methods capable of using underlying biological structure, but the development of such models is hampered by the lack of published benchmark tasks and well characterized baselines. In this work, we establish such benchmarks and baselines by profiling many classifiers against biologically motivated tasks on two curated views of a large, public gene expression dataset (the LINCS corpus) and one privately produced dataset. We provide these two curated views of the public LINCS dataset and our benchmark tasks to enable direct comparisons to future methodological work and help spur deep learning method development on this modality. In addition to profiling a battery of traditional classifiers, including linear models, random forests, decision trees, K nearest neighbor (KNN) classifiers, and feed-forward artificial neural networks (FF-ANNs), we also test a method novel to this data modality: graph convolugtional neural networks (GCNNs), which allow us to incorporate prior biological domain knowledge. We find that GCNNs can be highly performant, with large datasets, whereas FF-ANNs consistently perform well. Non-neural classifiers are dominated by linear models and KNN classifiers.
科研通智能强力驱动
Strongly Powered by AbleSci AI