Classification of DNA sequences using Bloom filters

Classification of DNA sequences using Bloom filters
复制标题

DOI:
10.1093/bioinformatics/btq230
复制
发表时间:
2010-07-01
期刊:
影响因子:
5.8
通讯作者:
Lundeberg, Joakim
Lundeberg, Joakim
中科院分区:
生物学3区
文献类型:
--
作者:
Stranneheim, Henrik;Kaller, Max;Lundeberg, Joakim

文献摘要

被引文献

相似文献

动机:新一代测序技术产生越来越复杂的数据集,需要新的高效和专业的序列分析算法。通常情况下,它是唯一的“新的”序列在一个复杂的数据集,是感兴趣的和多余的序列需要被removed.Results:一种新的算法,快速和准确的序列分类(FACS),介绍,可以准确和快速地分类序列属于或不属于一个参考序列。首先使用合成宏基因组数据集优化和验证FACS。然后使用实验宏基因组数据集来显示FACS实现与BLAT和SSAHA2相当的准确性,但在分类序列方面至少快21倍。
Motivation: New generation sequencing technologies producing increasingly complex datasets demand new efficient and specialized sequence analysis algorithms. Often, it is only the 'novel' sequences in a complex dataset that are of interest and the superfluous sequences need to be removed.Results: A novel algorithm, fast and accurate classification of sequences (FACSs), is introduced that can accurately and rapidly classify sequences as belonging or not belonging to a reference sequence. FACS was first optimized and validated using a synthetic metagenome dataset. An experimental metagenome dataset was then used to show that FACS achieves comparable accuracy as BLAT and SSAHA2 but is at least 21 times faster in classifying sequences.