聚类分析
相似性(几何)
模式识别(心理学)
病理学
人工智能
计算机科学
离群值
生物学数据
公制(单位)
轮廓
计算生物学
数据挖掘
机器学习
生物信息学
生物
医学
图像(数学)
精神科
运营管理
经济
作者
Lucia Prieto Santamaría,Eduardo P. García del Valle,Gerardo Lagunes García,Massimiliano Zanin,Alejandro Rodríguez‐González,Ernestina Menasalvas,Yuliana Pérez Gallardo,Gandhi Hernández-Chan
出处
期刊:
[Cold Spring Harbor Laboratory]
日期:2020-04-10
被引量:1
标识
DOI:10.1101/2020.04.10.035394
摘要
Abstract While classical disease nosology is based on phenotypical characteristics, the increasing availability of biological and molecular data is providing new understanding of diseases and their underlying relationships, that could lead to a more comprehensive paradigm for modern medicine. In the present work, similarities between diseases are used to study the generation of new possible disease nosologic models that include both phenotypical and biological information. To this aim, disease similarity is measured in terms of disease feature vectors, that stood for genes, proteins, metabolic pathways and PPIs in the case of biological similarity, and for symptoms in the case of phenotypical similarity. An improvement in similarity computation is proposed, considering weighted instead of Booleans feature vectors. Unsupervised learning methods were applied to these data, specifically, density-based DBSCAN clustering algorithm. As evaluation metric silhouette coefficient was chosen, even though the number of clusters and the number of outliers were also considered. As a results validation, a comparison with randomly distributed data was performed. Results suggest that weighted biological similarities based on proteins, and computed according to cosine index, may provide a good starting point to rearrange disease taxonomy and nosology.
科研通智能强力驱动
Strongly Powered by AbleSci AI