过度拟合
计算机科学
分类器(UML)
加权
模式识别(心理学)
卷积神经网络
人工智能
基本事实
数据挖掘
机器学习
噪声数据
人工神经网络
医学
放射科
作者
Fang Jing,Shao‐Wu Zhang,Shihua Zhang
出处
期刊:Methods
[Elsevier BV]
日期:2022-04-21
卷期号:203: 207-213
被引量:6
标识
DOI:10.1016/j.ymeth.2022.04.010
摘要
With the accumulation of ChIP-seq data, convolution neural network (CNN)-based methods have been proposed for predicting transcription factor binding sites (TFBSs). However, biological experimental data are noisy, and are often treated as ground truth for both training and testing. Particularly, existing classification methods ignore the false positive and false negative which are caused by the error in the peak calling stage, and therefore, they can easily overfit to biased training data. It leads to inaccurate identification and inability to reveal the rules of governing protein-DNA binding. To address this issue, we proposed a meta learning-based CNN method (namely TFBS_MLCNN or MLCNN for short) for suppressing the influence of noisy labels data and accurately recognizing TFBSs from ChIP-seq data. Guided by a small amount of unbiased meta-data, MLCNN can adaptively learn an explicit weighting function from ChIP-seq data and update the parameter of classifier simultaneously. The weighting function overcomes the influence of biased training data on classifier by assigning a weight to each sample according to its training loss. The experimental results on 424 ChIP-seq datasets show that MLCNN not only outperforms other existing state-of-the-art CNN methods, but can also detect noisy samples which are given the small weights to suppress them. The suppression ability to the noisy samples can be revealed through the visualization of samples' weights. Several case studies demonstrate that MLCNN has superior performance to others.
科研通智能强力驱动
Strongly Powered by AbleSci AI