化学空间
计算机科学
人工智能
范式转换
自编码
深度学习
生成语法
财产(哲学)
机器学习
相似性(几何)
人工神经网络
药效团
特征(语言学)
语言模型
特征学习
认知科学
理论计算机科学
代表(政治)
变压器
对抗制
空格(标点符号)
循环神经网络
学习迁移
生成模型
化学信息学
自然语言处理
作者
Arkaprava Banerjee,Supratik Kar,Kunal Roy,Grace Patlewicz,Imran Shah,Panagiotis G. Karamertzanis,Giuseppina Gini,Emilio Benfenati,Arkaprava Banerjee,Supratik Kar,Kunal Roy,Grace Patlewicz,Imran Shah,Panagiotis G. Karamertzanis,Giuseppina Gini,Emilio Benfenati
摘要
ABSTRACT This review provides a comprehensive overview of the paradigm shift for computer‐aided molecular design and property predictions from similarity‐based modeling, including quantitative structure–activity/property relationship (QSAR/QSPR), read‐across, read‐across structure–activity relationship (RASAR), and pharmacophore mapping to sequence‐based chemical language models (CLMs) using deep learning techniques. Starting with multiple methods of chemical structure and latent chemical space representations and touching the molecular descriptor‐ and fingerprint‐based classical type modeling, this review introduces string‐based deep learning models involving techniques like recurrent neural networks (RNNs) with long short‐term memory (LSTM) and other architectures such as variational autoencoder (VAE), attention models, and generative adversarial networks (GANs). The basics of more efficient transformer models are also discussed. The problem‐solving of training with scarce data using transfer learning, data augmentation, and natural‐product‐inspired training is analyzed. The applications of CLMs in the de novo design of small molecules of medicinal interest, enzymes, peptides, and multitask agents, the predictions of properties of drug candidates, and activity cliffs are presented. The applications of CLMs in materials science and predictive toxicology are also mentioned. We discuss the limitations of feature‐based modeling approaches confined to a restricted feature space. In contrast, CLMs lack specific insights into aspects like SARs, bioisosteric replacements, synthesizability, and so forth, which collectively hinder their regulatory acceptance and acceptance by synthetic chemists. This review concludes that cheminformaticians need to utilize two complementary approaches, where factors like simplicity, reproducibility, and regulatory acceptability may prompt the use of feature‐based approaches while aiming for higher accuracy and generating novel molecules may drive toward adopting CLMs. This article is categorized under: Data Science > Chemoinformatics Structure and Mechanism > Computational Biochemistry and Biophysics Software > Molecular Modeling
科研通智能强力驱动
Strongly Powered by AbleSci AI