计算机科学
缺少数据
数据挖掘
模块化设计
鉴定(生物学)
概率逻辑
机器学习
工作流程
贝叶斯概率
判别式
稳健性(进化)
稀缺
人工智能
不完美的
特征(语言学)
班级(哲学)
样品(材料)
初始化
预测建模
钥匙(锁)
离群值
点(几何)
集合预报
点估计
贝叶斯定理
统计模型
数据点
一致性(知识库)
数据建模
残余物
作者
Hanle Lin,Hao Wen,Zheng Ma,Y Q Zhao,Qianwen Zheng,Chuyuan He,Lekang Cui,Yayun Zhang
标识
DOI:10.1021/acs.est.6c05983
摘要
Environmental quantitative structure-activity relationship (QSAR) modeling is frequently constrained by small sample sizes, missing values, and end point imbalance. To address these limitations, we propose a modular framework based on the TabPFN family that integrates zero-/few-shot prediction, missing robust probabilistic inference, and generative augmentation. Validated across hydroxyl radical rate prediction, singlet oxygen kinetics, and toxicity classification tasks, the framework demonstrates strong data efficiency. Data utility curves combined with a dynamic efficient-window identification strategy provide data-driven guidance for sample size selection. The framework also maintains stable predictive performance and preserves key feature identification under moderate missingness. Under extreme data scarcity (<100 samples) and severe class imbalance, the generative module improves certain conventional classifiers, outperforming GMM-based augmentation and achieving performance comparable to SMOTE while providing limited benefit for TabPFN itself. SHAP and PDP analyses further reveal chemically meaningful trends consistent with current mechanistic knowledge. Overall, this study provides a flexible QSAR workflow for reliable, interpretable, and cost-aware modeling under real-world data constraints.
科研通智能强力驱动
Strongly Powered by AbleSci AI