串联(数学)
异方差
统计
校准
数学
重采样
计算机科学
随机性
相关性
分布(数学)
缺少数据
时间点
酶动力学
点(几何)
计量经济学
人工智能
色散(光学)
不相关
作者
Xinran Wang,Wuruiyang Li,Shuaiwen Ding,Yi Deng,Haijuan Zhang,Hongyu Li,Yang Li
标识
DOI:10.1021/acs.jcim.6c02075
摘要
Abstract Public enzyme-kinetics records have labels that are incomplete across tasks, assay conditions that vary across sources, and structure-associated evidence that is not uniformly available. Consequently, naive multimodal concatenation is limited, and scores from random splits are weak proxies for deployment performance. We present KinEAGER, an evidence-aware multitask framework that jointly predicts kcat and Km and analytically derives kcat/Km under missing modalities and distribution shift. KinEAGER combines frozen ESM-2 and MolT5 encoders, low-rank adapters (LoRA), cross-modal interaction layers, availability-aware soft gating, domain-by-task heteroscedastic training, and a kcat specialist engaged according to distance from the training distribution for out-of-distribution (OOD) queries; three independently trained models form a deep ensemble that supplies sample-level uncertainty. On a grouped enzyme–substrate split, kcat reaches a MAE of 0.719 (R2 = 0.542) and Km reaches a MAE of 0.722 (R2 = 0.737) in log10 units, with a macro-average MAE of 0.817. Under sequence-cluster OOD evaluation, KinEAGER achieves the lowest MAE among the locally evaluated models and approaches the performance reported for CatPred. Three-run ablations show a consistent pattern: the Full configuration gives the best overall performance, structure randomization causes the largest degradation in point prediction, domain-by-task uncertainty primarily improves NLL and coverage, and either physicochemical or structural evidence alone is weaker than the Full configuration. Prediction disagreement among the three independent models is positively correlated with absolute error (macro-average Spearman correlation 0.144). When ensemble disagreement is used as a split-conformal normalizer, validation-calibrated sample-adaptive intervals achieve empirical coverage of 0.903/0.951 at the nominal 90%/95% levels (Cov90/Cov95). Calibration degrades on new data sources, while a small target-source calibration set can restore coverage. A Nitrocefin assay on four beta-lactamases reproduces the predicted HIGH/LOW ranking, with KPC-family 0G8 exceeding the CTX-M-1 reference.
科研通智能强力驱动
Strongly Powered by AbleSci AI