相似性(几何)
树(集合论)
公制(单位)
计算机科学
数据挖掘
排名(信息检索)
系统发育树
聚类分析
独立性(概率论)
概率逻辑
度量(数据仓库)
人工智能
数学
理论计算机科学
机器学习
情报检索
统计
组合数学
生物化学
运营管理
化学
经济
图像(数学)
基因
出处
期刊:Bioinformatics
[Oxford University Press]
日期:2020-06-26
卷期号:36 (20): 5007-5013
被引量:139
标识
DOI:10.1093/bioinformatics/btaa614
摘要
Abstract Motivation The Robinson–Foulds (RF) metric is widely used by biologists, linguists and chemists to quantify similarity between pairs of phylogenetic trees. The measure tallies the number of bipartition splits that occur in both trees—but this conservative approach ignores potential similarities between almost-identical splits, with undesirable consequences. ‘Generalized’ RF metrics address this shortcoming by pairing splits in one tree with similar splits in the other. Each pair is assigned a similarity score, the sum of which enumerates the similarity between two trees. The challenge lies in quantifying split similarity: existing definitions lack a principled statistical underpinning, resulting in misleading tree distances that are difficult to interpret. Here, I propose probabilistic measures of split similarity, which allow tree similarity to be measured in natural units (bits). Results My new information-theoretic metrics outperform alternative measures of tree similarity when evaluated against a broad suite of criteria, even though they do not account for the non-independence of splits within a single tree. Mutual clustering information exhibits none of the undesirable properties that characterize other tree comparison metrics, and should be preferred to the RF metric. Availability and implementation The methods discussed in this article are implemented in the R package ‘TreeDist’, archived at https://dx.doi.org/10.5281/zenodo.3528123. Supplementary information Supplementary data are available at Bioinformatics online.
科研通智能强力驱动
Strongly Powered by AbleSci AI