财产(哲学)
计算机科学
人工智能
图形
机器学习
平滑度
悬崖
回归
过程(计算)
基础(线性代数)
训练集
数据挖掘
模式识别(心理学)
指纹(计算)
一致性(知识库)
理论计算机科学
算法
均方预测误差
透视图(图形)
高斯过程
作者
Chao Cui,Xiaorui Su,Zaixi Zhang,Alejandro Velez-Arce,Jianmin Wang,Xiang Cheng Shi,Yan Bing Zhang,J R Wu,Yu Zong Chen,Marinka Žitnik,Wan Xiang Shen
标识
DOI:10.26434/chemrxiv-2023-5cz7s/v4
摘要
Accurately predicting molecular activity is hindered by activity cliffs (ACs), which are sharp potency changes between highly similar compounds that distort the smoothness assumed by modern quantitative structure–activity relationship (SAR) models and graph neural networks (GNNs). Here we introduce AC awareness (ACA), an inductive bias that reshapes GNN latent spaces to account for these discontinuities. Implemented through an ACA loss combining regression with softmargin triplet contrastive learning, the method dynamically mines high-value activity-cliff triplets during training and corrects inconsistent neighbourhoods in latent space. This process yields progressively fewer cliff violations, more coherent activity gradients, and reduced label incoherence across diverse chemical spaces. Evaluated on 52 datasets spanning low-sample narrow-scaffold series, large mixed-scaffold benchmarks, matched-pair cliff classification, and absorption, distribution, metabolism, excretion and toxicity (ADMET) property prediction, ACA consistently improves predictive accuracy over strong extended-connectivity fingerprint (ECFP) and GNN baselines. The approach generalizes across multiple GNN backbones and remains robust under fixed hyperparameters. These results suggest that ACA provides a principled strategy for enhancing molecular property prediction by aligning latent representations with the nonadditive behaviour underlying ACs.
科研通智能强力驱动
Strongly Powered by AbleSci AI