Blobology: exploring raw genome data for contaminants, symbionts and parasites using taxon-annotated GC-coverage plots.

Blobology: exploring raw genome data for contaminants, symbionts and parasites using taxon-annotated GC-coverage plots.
复制标题

DOI:
10.3389/fgene.2013.00237
复制
发表时间:
2013
影响因子:
3.7
通讯作者:
Blaxter M
Blaxter M
中科院分区:
生物学3区
文献类型:
--
作者:
Kumar S;Jones M;Koutsovoulos G;Clarke M;Blaxter M

文献摘要

参考文献

被引文献

相似文献

为靶真核物种的从头基因组组装项目生成原始数据相对容易。大规模数据的民主化使得许多研究团队能够计划组装非模式生物的基因组。这些新的基因组靶标与传统的、近交的、实验室饲养的模式生物非常不同。它们通常很小,不能脱离其环境----无论是摄入的食物、周围的寄生虫宿主生物体,还是附着在采样个体上或采样个体内的共生和共生生物体。制备来自单一物种的纯DNA在技术上是不可能的,但组装混合生物体DNA可能很困难,因为大多数基因组组装者在面对不同化学计量的多个基因组时表现不佳。这类问题在宏基因组数据集中很常见,这些数据集故意试图捕获环境中存在的所有基因组,但复制子组装通常不是此类程序的目标。在这里,我们提出了一种方法来提取,从混合的DNA序列数据,子集对应于单一物种的基因组,从而提高基因组组装。我们使用数字(GC碱基和读段覆盖率的比例)和生物(注释数据库中的最佳匹配序列)指标来帮助将草案组装重叠群和有助于这些重叠群的读段划分为不同的箱,然后可以通过使用分类单位注释的GC覆盖率图(TAGC图)进行严格的优化组装。我们还介绍了Blobsporer,这是一种帮助从TAGC注释数据中探索和选择子集的工具。以这种方式划分数据可以拯救组装不良的基因组,并揭示真核基因组计划中意想不到的共生体和共生体。TAGC plot pipeline脚本可从https://github.com/blaxterlab/blobology获得,Blobsporer工具可从https://github.com/mojones/Blobsplorer获得。
Generating the raw data for a de novo genome assembly project for a target eukaryotic species is relatively easy. This democratization of access to large-scale data has allowed many research teams to plan to assemble the genomes of non-model organisms. These new genome targets are very different from the traditional, inbred, laboratory-reared model organisms. They are often small, and cannot be isolated free of their environment – whether ingested food, the surrounding host organism of parasites, or commensal and symbiotic organisms attached to or within the individuals sampled. Preparation of pure DNA originating from a single species can be technically impossible, but assembly of mixed-organism DNA can be difficult, as most genome assemblers perform poorly when faced with multiple genomes in different stoichiometries. This class of problem is common in metagenomic datasets that deliberately try to capture all the genomes present in an environment, but replicon assembly is not often the goal of such programs. Here we present an approach to extracting, from mixed DNA sequence data, subsets that correspond to single species’ genomes and thus improving genome assembly. We use both numerical (proportion of GC bases and read coverage) and biological (best-matching sequence in annotated databases) indicators to aid partitioning of draft assembly contigs, and the reads that contribute to those contigs, into distinct bins that can then be subjected to rigorous, optimized assembly, through the use of taxon-annotated GC-coverage plots (TAGC plots). We also present Blobsplorer, a tool that aids exploration and selection of subsets from TAGC-annotated data. Partitioning the data in this way can rescue poorly assembled genomes, and reveal unexpected symbionts and commensals in eukaryotic genome projects. The TAGC plot pipeline script is available from https://github.com/blaxterlab/blobology, and the Blobsplorer tool from https://github.com/mojones/Blobsplorer.
DOI: 10.1186/1471-2164-14-87
发表时间: 2013-02-08
期刊: BMC genomics
影响因子: 4.4
作者:
Heitlinger E;Bridgett S;Montazam A;Taraschewski H;Blaxter M
通讯作者: Blaxter M
DOI: 10.4161/worm.19046
发表时间: 2012-01-01
期刊: Worm
影响因子: --
作者:
Kumar S;Koutsovoulos G;Kaur G;Blaxter M
通讯作者: Blaxter M
DOI: 10.1038/nmeth.1923
发表时间: 2012-03-04
期刊: NATURE METHODS
影响因子: 48
作者:
Langmead, Ben;Salzberg, Steven L.
通讯作者: Salzberg, Steven L.
DOI: 10.1186/gb-2013-14-8-r89
发表时间: 2013-08-28
期刊: Genome biology
影响因子: 12.3
作者:
Schwarz EM;Korhonen PK;Campbell BE;Young ND;Jex AR;Jabbar A;Hall RS;Mondal A;Howe AC;Pell J;Hofmann A;Boag PR;Zhu XQ;Gregory T;Loukas A;Williams BA;Antoshechkin I;Brown C;Sternberg PW;Gasser RB
通讯作者: Gasser RB
DOI: 10.1186/1471-2105-6-31
发表时间: 2005-02-15
期刊: BMC bioinformatics
影响因子: 3
作者:
Slater GS;Birney E
通讯作者: Birney E