熔点
符号
计算机科学
自然(考古学)
分子
有机分子
点(几何)
人工智能
化学
自然语言处理
有机化学
语言学
数学
哲学
地质学
古生物学
几何学
作者
Weiming Mi,Huijun Chen,Donghua Zhu,Tao Zhang,Feng Qian
摘要
Establishing quantitative structure-property relationships for the rational design of small molecule drugs at the early discovery stage is highly desirable. Using natural language processing (NLP), we proposed a machine learning model to process the line notation of small organic molecules, allowing the prediction of their melting points. The model prediction accuracy benefits from training upon different canonicalized SMILES forms of the same molecules and does not decrease with increasing size, complexity, and structural flexibility. When a combination of two different canonicalized SMILES forms is used to train the model, the prediction accuracy improves. Largely distinguished from the previous fragment-based or descriptor-based models, the prediction accuracy of this NLP-based model does not decrease with increasing size, complexity, and structural flexibility of molecules. By representing the chemical structure as a natural language, this NLP-based model offers a potential tool for quantitative structure-property prediction for drug discovery and development.
科研通智能强力驱动
Strongly Powered by AbleSci AI