SEED: efficient clustering of next-generation sequences

SEED: efficient clustering of next-generation sequences
复制标题

DOI:
10.1093/bioinformatics/btr447
复制
发表时间:
2011-09-15
期刊:
影响因子:
5.8
通讯作者:
Girke, Thomas
Girke, Thomas
中科院分区:
生物学3区
文献类型:
--
作者:
Bao, Ergude;Jiang, Tao;Girke, Thomas

文献摘要

被引文献

相似文献

动机:下一代序列的相似性聚类是研究DNA/RNA分子群体大小和减少下一代序列数据冗余的重要计算问题。目前,大多数序列聚类算法是有限的,他们的速度和可扩展性,从而不能处理数据与数以千万计的reads.Results:在这里,我们介绍了种子-一个高效的算法聚类非常大的NGS集。它将序列连接成簇,这些簇可以与它们的虚拟中心相差多达三个错配和三个突出残基。它是基于一种改进的间隔种子方法,称为块间隔种子。它的聚类组件通过首先识别虚拟中心序列,然后找到所有满足相似性参数的相邻序列来对哈希表进行操作。SEED可以在< 4小时内对1亿个短读取序列进行集群,并具有线性时间和内存性能。当使用SEED作为基因组/转录组组装数据的预处理工具时,它能够将本研究中使用的数据集的Velvet/Oasis组装器的时间和内存需求分别减少60-85%和21- 41%。此外,组装物含有比未预处理的数据更长的重叠群,如大12-27%的N50值所示。与其他聚类工具相比,SEED在生成与真实聚类结果相似的NGS数据聚类方面表现出最佳性能,时间性能提高了2到10倍。虽然SEED的大多数实用程序都属于NGS数据的预处理领域,但我们的测试也证明了它作为独立工具的效率,用于从未测序的生物体中发现NGS数据中的小RNA序列簇。
Motivation: Similarity clustering of next generation sequences (NGS) is an important computational problem to study the population sizes of DNA/RNA molecules and to reduce the redundancies in NGS data. Currently, most sequence clustering algorithms are limited by their speed and scalability, and thus cannot handle data with tens of millions of reads.Results: Here, we introduce SEED-an efficient algorithm for clustering very large NGS sets. It joins sequences into clusters that can differ by up to three mismatches and three overhanging residues from their virtual center. It is based on a modified spaced seed method, called block spaced seeds. Its clustering component operates on the hash tables by first identifying virtual center sequences and then finding all their neighboring sequences that meet the similarity parameters. SEED can cluster 100 million short read sequences in < 4 h with a linear time and memory performance. When using SEED as a preprocessing tool on genome/transcriptome assembly data, it was able to reduce the time and memory requirements of the Velvet/Oasis assembler for the datasets used in this study by 60-85% and 21-41%, respectively. In addition, the assemblies contained longer contigs than non-preprocessed data as indicated by 12-27% larger N50 values. Compared with other clustering tools, SEED showed the best performance in generating clusters of NGS data similar to true cluster results with a 2- to 10-fold better time performance. While most of SEED's utilities fall into the preprocessing area of NGS data, our tests also demonstrate its efficiency as stand-alone tool for discovering clusters of small RNA sequences in NGS data from unsequenced organisms.