Exploring Pandora's Box: Potential and Pitfalls of Low Coverage Genome Surveys for Evolutionary Biology

Exploring Pandora's Box: Potential and Pitfalls of Low Coverage Genome Surveys for Evolutionary Biology
复制标题

DOI:
10.1371/journal.pone.0049202
复制
发表时间:
2012-11-21
期刊:
影响因子:
3.7
通讯作者:
Sands, Chester J.
Sands, Chester J.
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Leese, Florian;Brand, Philipp;Sands, Chester J.

文献摘要

被引文献

相似文献

高通量测序技术正在给基因研究带来革命性的变化。随着“机器的兴起”,即使是未知的基因组,也可以在很短的时间内以合理的成本获得基因组序列。这使得进化生物学家研究遗传上未被探索的物种时,只需对一小部分基因组进行测序,就能确定分子标记或感兴趣的基因组区域(如微卫星和小卫星、线粒体和核基因)。然而,当使用这些来自非模式物种的数据集时,来自非目标污染物物种(如细菌、病毒、真菌或其他真核生物)的DNA可能会使结果的解释复杂化。本研究分析了来自4个主要进化谱系的水生非模式类群的14个基因组焦磷酸测序文库。我们量化了合适的微卫星和小卫星、线粒体基因组、已知核基因和转座因子的数量,并使用生物信息学方法搜索来自各种来源的污染。结果表明,在所有覆盖范围约为0.02-25%的序列文库中,可以从不同的KEGG(京都基因与基因组百科全书)途径中识别出许多合适的微卫星和小卫星、线粒体基因序列和核基因。这些可以作为系统发育和群体遗传分析的标记。我们研究的一个中心发现是,由于非目标DNA或移动元素,几个基因组文库遭受了不同的偏差。特别是,病毒、细菌或真核生物内共生体对分析的一些文库贡献很大(高达10%)。如果不是这样,从非模式生物的高通量测序数据中开发的遗传标记可能会影响进化研究或在实验测试中完全失败。总之,我们的研究证明了低覆盖率基因组调查序列的巨大潜力,并提出了生物信息学分析工作流程。结果还建议在开发标记之前对有问题的序列和非目标基因组序列进行更复杂的过滤。
High throughput sequencing technologies are revolutionizing genetic research. With this "rise of the machines'', genomic sequences can be obtained even for unknown genomes within a short time and for reasonable costs. This has enabled evolutionary biologists studying genetically unexplored species to identify molecular markers or genomic regions of interest (e. g. micro- and minisatellites, mitochondrial and nuclear genes) by sequencing only a fraction of the genome. However, when using such datasets from non-model species, it is possible that DNA from non-target contaminant species such as bacteria, viruses, fungi, or other eukaryotic organisms may complicate the interpretation of the results. In this study we analysed 14 genomic pyrosequencing libraries of aquatic non-model taxa from four major evolutionary lineages. We quantified the amount of suitable micro-and minisatellites, mitochondrial genomes, known nuclear genes and transposable elements and searched for contamination from various sources using bioinformatic approaches. Our results show that in all sequence libraries with estimated coverage of about 0.02-25%, many appropriate micro-and minisatellites, mitochondrial gene sequences and nuclear genes from different KEGG (Kyoto Encyclopedia of Genes and Genomes) pathways could be identified and characterized. These can serve as markers for phylogenetic and population genetic analyses. A central finding of our study is that several genomic libraries suffered from different biases owing to non-target DNA or mobile elements. In particular, viruses, bacteria or eukaryote endosymbionts contributed significantly (up to 10%) to some of the libraries analysed. If not identified as such, genetic markers developed from high-throughput sequencing data for non-model organisms may bias evolutionary studies or fail completely in experimental tests. In conclusion, our study demonstrates the enormous potential of low-coverage genome survey sequences and suggests bioinformatic analysis workflows. The results also advise a more sophisticated filtering for problematic sequences and non-target genome sequences prior to developing markers.