已入深夜,您辛苦了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!祝你早点完成任务,早点休息,好梦!

A comprehensive pipeline to integrate preprocessing and machine learning techniques for accurate classification in Raman spectroscopy

计算机科学 预处理器 人工智能 过度拟合 数据预处理 稳健性(进化) 机器学习 降维 特征提取 数据挖掘 模式识别(心理学) 人工神经网络 生物化学 基因 化学
作者
Simone Innocente,Siddra Maryam,Stefan Andersson‐Engels,Katarzyna Komolibus,Rekha Gautam,Andrea Visentin
标识
DOI:10.1117/12.3017024
摘要

Raman spectroscopy, a non-invasive analytical method, offers insights into molecular structures and interactions in various liquid and solid samples with applications ranging from material science, and chemical analysis to medical diagnostics. Preprocessing of Raman spectra is vital to remove interferences like background signals and calibration errors, ensuring precise data extraction. Artificial intelligence, particularly machine learning (ML), aids in extracting valuable information from complex datasets. However, effective data preprocessing proves to be crucial as it can influence model robustness. This study addresses the integration of preprocessing and ML algorithms, often treated as distinct identities despite their intrinsic interconnection, in Raman spectra of blood samples from patients suffering from ovarian cancer. Optimal preprocessing configuration may not always be evident due to the complexity of spectral data. There are numerous options available for background corrections, normalization, outlier removal, noise filtering, and dimension reduction algorithms for Raman spectra. Moreover, hyperparameter tuning is required to detect the best choices for the preprocessing steps. In this work, we present a pipeline to co-optimize preprocessing techniques and ML classification methods to promote objective selection and minimize processing time. In our approach, preprocessing methods are not chosen arbitrarily but rather systematically evaluated to enhance the robustness of the models. These criteria focus on ensuring that the model performs well not only on the training data but also on unseen data, thus reducing the risk of overfitting and improving the generalization capability of the model. This systematic approach would reduce the time for new studies by detecting the most suitable preprocessing steps and hyperparameters needed and building a robust model for the task.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
tt完成签到,获得积分10
刚刚
Lucas应助悦耳青梦采纳,获得10
2秒前
浮沉发布了新的文献求助10
3秒前
day1发布了新的文献求助20
4秒前
JamesPei应助杭黎昕采纳,获得10
6秒前
6秒前
脑洞疼应助杭黎昕采纳,获得10
6秒前
桐桐应助杭黎昕采纳,获得10
6秒前
李健的小迷弟应助杭黎昕采纳,获得10
6秒前
脑洞疼应助杭黎昕采纳,获得10
6秒前
公冶凡波完成签到,获得积分10
6秒前
可爱的函函应助杭黎昕采纳,获得10
6秒前
小马甲应助杭黎昕采纳,获得10
6秒前
JamesPei应助杭黎昕采纳,获得10
7秒前
Ava应助杭黎昕采纳,获得10
7秒前
搜集达人应助杭黎昕采纳,获得10
7秒前
风趣小松鼠完成签到 ,获得积分10
9秒前
李健的粉丝团团长应助Leo采纳,获得10
9秒前
9秒前
11秒前
12秒前
12秒前
12秒前
mmyhn应助day1采纳,获得20
13秒前
852应助杭黎昕采纳,获得10
15秒前
悦耳青梦发布了新的文献求助10
15秒前
wanci应助杭黎昕采纳,获得10
15秒前
在水一方应助杭黎昕采纳,获得10
15秒前
深情安青应助杭黎昕采纳,获得10
15秒前
汉堡包应助杭黎昕采纳,获得10
15秒前
CipherSage应助杭黎昕采纳,获得10
16秒前
Owen应助杭黎昕采纳,获得10
16秒前
NexusExplorer应助杭黎昕采纳,获得10
16秒前
16秒前
NexusExplorer应助杭黎昕采纳,获得10
16秒前
Lucas应助杭黎昕采纳,获得10
16秒前
17秒前
Chris发布了新的文献求助10
18秒前
20秒前
德芙发布了新的文献求助10
21秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Römisch-Germanische Forschungen 1000
APA handbook of comparative psychology: Basic concepts, methods, neural substrate, and behavior 1000
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The fast track to determining transfer functions of linear circuits: The student guide 500
Electric machines: theory, operating applications, and controls 500
The Analytical and Numerical Solution of Electric and Magnetic Fields 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7604398
求助须知:如何正确求助?哪些是违规求助? 9180283
关于积分的说明 19661255
捐赠科研通 7179519
什么是DOI,文献DOI怎么找? 3269377
关于科研通互助平台的介绍 2433373
邀请新用户注册赠送积分活动 2263437