Artificial and natural duplicates in pyrosequencing reads of metagenomic data.

Artificial and natural duplicates in pyrosequencing reads of metagenomic data.
复制标题

DOI:
10.1186/1471-2105-11-187
复制
发表时间:
2010-04-13
期刊:
影响因子:
3
通讯作者:
Li W
Li W
中科院分区:
生物学4区
文献类型:
--
作者:
Niu B;Fu L;Sun S;Li W

文献摘要

参考文献

被引文献

相似文献

焦磷酸测序读段中的人工重复可能会导致在宏基因组研究中对物种和基因丰度的错误解读。在许多宏基因组项目中,重复读段被过滤掉。然而,由于在一次焦磷酸测序运行中观察到的重复读段也包括自然(非人工)重复,简单地去除所有重复可能也会导致对与自然重复相关的丰度的低估。 我们实施了一种从焦磷酸测序读段中识别完全相同和近乎相同重复的方法。该方法进行所有对所有的序列比较,并使用从我们之前的序列聚类方法cd - hit修改而来的算法将重复聚类成组。这种方法可以在约10分钟内处理一个典型的数据集;它还为每组重复提供一个共有序列。我们将这种方法应用于39个基因组项目和10个使用焦磷酸测序技术的宏基因组项目的原始读段。我们比较了通过我们的方法识别出的重复以及通过独立模拟产生的自然重复的出现情况。我们观察到,包括人工和自然重复在内的重复占读段的4 - 44%。自然重复的数量与样本的读段密度(读段数量除以基因组大小)高度相关。对于缺乏优势物种的高复杂性宏基因组样本,自然重复仅占所有重复的<1%。但对于一些其他样本,如转录组样本,观察到的大多数重复可能是自然重复。 我们的方法可从http://cd - hit.org获取,作为一个可下载的程序和一个网络服务器。不仅从宏基因组数据集中识别重复,而且区分它们是人工重复还是自然重复是很重要的。我们提供了一种根据用户定义的样本类型估计自然重复数量的工具,这样用户就可以决定在他们的项目中是保留还是去除重复。
Artificial duplicates from pyrosequencing reads may lead to incorrect interpretation of the abundance of species and genes in metagenomic studies. Duplicated reads were filtered out in many metagenomic projects. However, since the duplicated reads observed in a pyrosequencing run also include natural (non-artificial) duplicates, simply removing all duplicates may also cause underestimation of abundance associated with natural duplicates. We implemented a method for identification of exact and nearly identical duplicates from pyrosequencing reads. This method performs an all-against-all sequence comparison and clusters the duplicates into groups using an algorithm modified from our previous sequence clustering method cd-hit. This method can process a typical dataset in ~10 minutes; it also provides a consensus sequence for each group of duplicates. We applied this method to the underlying raw reads of 39 genomic projects and 10 metagenomic projects that utilized pyrosequencing technique. We compared the occurrences of the duplicates identified by our method and the natural duplicates made by independent simulations. We observed that the duplicates, including both artificial and natural duplicates, make up 4-44% of reads. The number of natural duplicates highly correlates with the samples' read density (number of reads divided by genome size). For high-complexity metagenomic samples lacking dominant species, natural duplicates only make up <1% of all duplicates. But for some other samples like transcriptomic samples, majority of the observed duplicates might be natural duplicates. Our method is available from http://cd-hit.org as a downloadable program and a web server. It is important not only to identify the duplicates from metagenomic datasets but also to distinguish whether they are artificial or natural duplicates. We provide a tool to estimate the number of natural duplicates according to user-defined sample types, so users can decide whether to retain or remove duplicates in their projects.
DOI: 10.1126/science.1093857
发表时间: 2004-04-02
期刊: SCIENCE
影响因子: 56.9
作者:
Venter, JC;Remington, K;Smith, HO
通讯作者: Smith, HO
DOI: 10.1186/gb-2007-8-7-r143
发表时间: 2007
期刊: Genome biology
影响因子: 12.3
作者:
Huse SM;Huber JA;Morrison HG;Sogin ML;Welch DM
通讯作者: Welch DM
DOI: 10.1038/nmeth1043
发表时间: 2007-06-01
期刊: NATURE METHODS
影响因子: 48
作者:
Mavromatis, Konstantinos;Ivanova, Natalia;Kyrpides, Nikos C.
通讯作者: Kyrpides, Nikos C.
DOI: 10.1093/bioinformatics/17.3.282
发表时间: 2001-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Li, WZ;Jaroszewski, L;Godzik, A
通讯作者: Godzik, A
DOI: 10.1126/science.1124234
发表时间: 2006-06-02
期刊: SCIENCE
影响因子: 56.9
作者:
Gill, Steven R.;Pop, Mihai;Nelson, Karen E.
通讯作者: Nelson, Karen E.