聚类分析
冗余(工程)
计算机科学
身份(音乐)
数据库
序列(生物学)
序列数据库
数据挖掘
人工智能
生物
遗传学
操作系统
声学
基因
物理
作者
Weizhong Li,Lukasz Jaroszewski,Adam Godzik
出处
期刊:Bioinformatics
[Oxford University Press]
日期:2002-01-01
卷期号:18 (1): 77-82
被引量:502
标识
DOI:10.1093/bioinformatics/18.1.77
摘要
Abstract Motivation: Sequence clustering replaces groups of similar sequences in a database with single representatives. Clustering large protein databases like the NCBI Non-Redundant database (NR) using even the best currently available clustering algorithms is very time-consuming and only practical at relatively high sequence identity thresholds. Our previous program, CD-HI, clustered NR at 90% identity in ∼1 h and at 75% identity in ∼1 day on a 1 GHz Linux PC (Li et al. , Bioinformatics, 17, 282, 2001); however even faster clustering speed is needed because the size of protein databases are rapidly growing and many applications desire a lower attainable thresholds. Results: For our previous algorithm (CD-HI), we have employed short-word filters to speed up the clustering. In this paper, we show that tolerating some redundancy makes for more efficient use of these short-word filters and increases the program’s speed 100 times. Our new program implements this technique and clusters NR at 70% identity within 2 h, and at 50% identity in ∼5 days. Although some redundancy is present after clustering, our new program’s results only differ from our previous program’s by less than 0.4%. Availability: The program and its previous version are available at http://bioinformatics.burnham-inst.org/cd-hi Contact: liwz@burnham-inst.org; adam@burnham-inst.org * To whom correspondence should be addressed.
科研通智能强力驱动
Strongly Powered by AbleSci AI