人工智能
计算机科学
机器学习
翻译(生物学)
模式识别(心理学)
机器翻译
数据集
水准点(测量)
集合(抽象数据类型)
鉴定(生物学)
谱线
训练集
支持向量机
反问题
融合
传感器融合
特征(语言学)
拉曼光谱
反向
算法
班级(哲学)
核磁共振谱数据库
特征提取
作者
Sheng Wu,Shiyang Ji,Mingzhi Yuan,Jun Jiang,Yi Luo,Wei Hu
标识
DOI:10.1021/acs.jpclett.5c02592
摘要
Accurate molecular structure recognition is vital in chemistry and drug discovery, yet single-spectrum methods often struggle with limited accuracy and robustness. Here, we present a multimodal machine learning framework based on our previously developed TranSpec to integrate IR, Raman, and NMR spectra for precise structure prediction. Taking a QM9S and QM9-NMR data set as the benchmark, we combined CNN and MLP in TranSpec to extract the local and global spectral features, and enhanced its translation accuracy using data augmentations (shifting, scaling, quantization) and ensemble voting. The results indicate that individual IR or Raman spectra substantially outperform NMR, with translation accuracies around 85%, compared to 30–50% for 13C and 1H NMR. We then employed data augmentation techniques and multimodal fusion to enhance the translation accuracy. Lastly, integrating the IR, Raman, 13C NMR, and 1H NMR spectra ultimately raised the Top-1 translation accuracy to 72.7% and the Top-3 accuracy to 97.7%. This study underscores the power of multimodal spectral fusion and establishes a new methodological benchmark for inverse molecular design in computational chemistry.
科研通智能强力驱动
Strongly Powered by AbleSci AI