已入深夜,您辛苦了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!祝你早点完成任务,早点休息,好梦!

AssayMatch: Learning To Select Data for Molecular Activity Models

计算机科学 机器学习 人工智能 训练集 数据挖掘 标记数据 数据质量 集合(抽象数据类型) 试验装置 数据集 数据建模 试验数据 镜像 噪音(视频) 合成数据 选择(遗传算法) 一致性(知识库) 选型 预测能力 实验数据 滤波器(信号处理) 判别式 数据收集 相似性(几何) 关系(数据库) 情报检索 质量(理念) 关系抽取 秩(图论) 灵活性(工程) 基线(sea)
作者
Vincent Fan,Regina Barzilay
出处
期刊:Journal of Chemical Information and Modeling [American Chemical Society]
卷期号:66 (7): 3540-3549
标识
DOI:10.1021/acs.jcim.5c02858
摘要

The performance of machine-learning models in drug discovery is highly dependent on the quality and consistency of the training data. Due to limitations in data set sizes, many models are trained by aggregating bioactivity data from diverse sources, including public databases such as ChEMBL. However, this approach often introduces significant noise due to variability in experimental protocols. We introduce AssayMatch, a framework for data selection that builds smaller, more homogeneous training sets attuned to the test set of interest. AssayMatch leverages data attribution methods to quantify the contribution of each training assay to the model's performance. These attribution scores are used to fine-tune language embeddings of text-based assay descriptions to capture not just semantic similarity but also the compatibility between assays. Unlike existing data attribution methods, our approach enables data selection for a test set with unknown labels, mirroring real-world drug discovery campaigns in which the activities of candidate molecules are not known in advance. At test time, embeddings fine-tuned with AssayMatch are used to rank all available training data. We demonstrate that models trained on data selected by AssayMatch are able to surpass the performance of the model trained on the complete data set, highlighting its ability to effectively filter out harmful or noisy experiments. We perform experiments on two common machine-learning architectures and see increased prediction capability over a strong language-only baseline for 8/12 model-target pairs. AssayMatch provides a data-driven mechanism to curate higher-quality data sets, reducing noise from incompatible experiments and improving the predictive power and data efficiency of models for drug discovery. AssayMatch is available at https://github.com/Ozymandias314/AssayMatch.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
威武苑睐发布了新的文献求助10
3秒前
zhenzheng完成签到 ,获得积分0
5秒前
7秒前
7秒前
7秒前
安静碧灵发布了新的文献求助10
11秒前
阳光不锈完成签到 ,获得积分10
11秒前
绘空事发布了新的文献求助10
12秒前
钱慧琳发布了新的文献求助10
13秒前
16秒前
香蕉觅云应助小映采纳,获得10
18秒前
22秒前
华仔应助顺顺尼采纳,获得10
23秒前
Ddd发布了新的文献求助10
24秒前
俊逸绮玉完成签到,获得积分20
25秒前
含蓄的芝麻完成签到,获得积分10
25秒前
木十四完成签到 ,获得积分10
29秒前
ABJ完成签到 ,获得积分10
29秒前
薛建伟完成签到,获得积分10
29秒前
31秒前
悦耳乘风完成签到,获得积分10
31秒前
香蕉觅云应助钱慧琳采纳,获得10
31秒前
32秒前
32秒前
小映发布了新的文献求助10
36秒前
36秒前
瘦瘦语兰发布了新的文献求助10
36秒前
39秒前
Ddd完成签到,获得积分20
41秒前
科研通AI6.4应助威武苑睐采纳,获得30
42秒前
琪琪发布了新的文献求助10
46秒前
52秒前
迷人绿蕊完成签到 ,获得积分10
53秒前
53秒前
Cher.完成签到,获得积分10
54秒前
david_x完成签到,获得积分10
55秒前
希望天下0贩的0应助1122采纳,获得10
55秒前
lqq完成签到 ,获得积分10
56秒前
56秒前
平淡的友儿完成签到 ,获得积分10
57秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
内視鏡的に摘除しえた十二指腸乳頭部腫瘍の2例 660
Cognitive Psychology in a Changing World 600
On nonlinear stability of contact discontinuities. In: Hyperbolic problems: theory, numerics, applications (Stony Brook, NY, 1994) 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
微电子器件实验教程 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7681220
求助须知:如何正确求助?哪些是违规求助? 9245502
关于积分的说明 19934825
捐赠科研通 7251720
什么是DOI,文献DOI怎么找? 3287810
关于科研通互助平台的介绍 2445511
邀请新用户注册赠送积分活动 2291387