Combinatorial pooling enables selective sequencing of the barley gene space.

Combinatorial pooling enables selective sequencing of the barley gene space.
复制标题

DOI:
10.1371/journal.pcbi.1003010
复制
发表时间:
2013-04
影响因子:
4.3
通讯作者:
Close TJ
Close TJ
中科院分区:
生物学2区
文献类型:
--
作者:
Lonardi S;Duma D;Alpert M;Cordero F;Beccuti M;Bhat PR;Wu Y;Ciardo G;Alsaihati B;Ma Y;Wanamaker S;Resnik J;Bozdag S;Luo MC;Close TJ

文献摘要

参考文献

被引文献

相似文献

对于绝大多数物种,包括许多经济或生态上重要的生物体,由于缺乏参考基因组序列,生物研究的进展受到阻碍。尽管测序技术最近取得了进展,但几个因素仍然限制了这种关键资源的可用性。与此同时,许多研究小组和国际财团已经建立了BAC文库和物理图谱,现在可以着手开发围绕固定在遗传图谱上的物理图谱组织的全基因组序列。我们提出了一个BAC的BAC测序协议,结合组合池设计和第二代测序技术,有效地接近从头选择性基因组测序。我们表明,组合池是一个具有成本效益的和实用的替代详尽的DNA条形码时,准备测序文库的数百或数千个DNA样品,如在这种情况下,基因承载的最小平铺路径BAC克隆。该方案的新奇取决于有效比较数亿个短读段并将其分配给正确的BAC克隆(去卷积)的计算能力,以便可以逐个克隆地进行组装。对水稻基因组模拟数据的实验结果表明,反卷积非常准确,得到的BAC组装体具有较高的质量。大麦基因组的基因丰富的子集的真实的数据的结果证实,解卷积是准确的,BAC组装具有良好的质量。虽然我们的方法不能提供一个全面的全基因组测序项目将实现的完整性水平,我们表明,它是相当成功的重建BAC内的基因序列。在植物如大麦的情况下,这种水平的序列知识足以支持关键的终点目标,如图位克隆和标记辅助育种。获得生物体完整基因组序列的问题已经通过全局蛮力方法(称为全基因组霰弹枪)或分而治之策略(称为逐个克隆)解决。这两种方法在成本、手工劳动以及处理测序错误和基因组高度重复区域的能力方面都有优点和缺点。随着第二代测序仪器的出现,全基因组鸟枪法已成为首选。然而,逐个克隆的策略对于大型复杂基因组仍然非常相关。事实上,几个研究小组和国际财团已经为许多经济或生态上重要的生物体制作了克隆文库和物理图谱,现在可以进行测序。在这篇手稿中,我们证明了这种方法在一个大的,非常重复的植物基因组的基因空间上的可行性。我们的方法的新奇在于,为了利用当前一代测序仪器的通量,我们使用特殊类型的“智能”池化设计来池化数百个克隆,该设计允许人们从池中的测序读段以高准确度建立源克隆。大量的模拟和实验结果支持我们的主张。
For the vast majority of species – including many economically or ecologically important organisms, progress in biological research is hampered due to the lack of a reference genome sequence. Despite recent advances in sequencing technologies, several factors still limit the availability of such a critical resource. At the same time, many research groups and international consortia have already produced BAC libraries and physical maps and now are in a position to proceed with the development of whole-genome sequences organized around a physical map anchored to a genetic map. We propose a BAC-by-BAC sequencing protocol that combines combinatorial pooling design and second-generation sequencing technology to efficiently approach denovo selective genome sequencing. We show that combinatorial pooling is a cost-effective and practical alternative to exhaustive DNA barcoding when preparing sequencing libraries for hundreds or thousands of DNA samples, such as in this case gene-bearing minimum-tiling-path BAC clones. The novelty of the protocol hinges on the computational ability to efficiently compare hundred millions of short reads and assign them to the correct BAC clones (deconvolution) so that the assembly can be carried out clone-by-clone. Experimental results on simulated data for the rice genome show that the deconvolution is very accurate, and the resulting BAC assemblies have high quality. Results on real data for a gene-rich subset of the barley genome confirm that the deconvolution is accurate and the BAC assemblies have good quality. While our method cannot provide the level of completeness that one would achieve with a comprehensive whole-genome sequencing project, we show that it is quite successful in reconstructing the gene sequences within BACs. In the case of plants such as barley, this level of sequence knowledge is sufficient to support critical end-point objectives such as map-based cloning and marker-assisted breeding. The problem of obtaining the full genomic sequence of an organism has been solved either via a global brute-force approach (called whole-genome shotgun) or by a divide-and-conquer strategy (called clone-by-clone). Both approaches have advantages and disadvantages in terms of cost, manual labor, and the ability to deal with sequencing errors and highly repetitive regions of the genome. With the advent of second-generation sequencing instruments, the whole-genome shotgun approach has been the preferred choice. The clone-by-clone strategy is, however, still very relevant for large complex genomes. In fact, several research groups and international consortia have produced clone libraries and physical maps for many economically or ecologically important organisms and now are in a position to proceed with sequencing. In this manuscript, we demonstrate the feasibility of this approach on the gene-space of a large, very repetitive plant genome. The novelty of our approach is that, in order to take advantage of the throughput of the current generation of sequencing instruments, we pool hundreds of clones using a special type of “smart” pooling design that allows one to establish with high accuracy the source clone from the sequenced reads in a pool. Extensive simulations and experimental results support our claims.
DOI: 10.1038/nmeth.1251
发表时间: 2008-10
期刊: NATURE METHODS
影响因子: 48
作者:
Craig, David W.;Pearson, John V.;Szelinger, Szabolcs;Sekar, Aswin;Redman, Margot;Corneveaux, Jason J.;Pawlowski, Traci L.;Laub, Trisha;Nunn, Gary;Stephan, Dietrich A.;Homer, Nils;Huentelman, Matthew J.
通讯作者: Huentelman, Matthew J.
DOI: 10.1101/gr.097261.109
发表时间: 2010-02-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Li, Ruiqiang;Zhu, Hongmei;Wang, Jun
通讯作者: Wang, Jun
DOI: 10.1109/tit.2009.2037043
发表时间: 2010-02
影响因子: 2.5
作者:
Erlich Y;Gordon A;Brand M;Hannon GJ;Mitra PP
通讯作者: Mitra PP
DOI: 10.1101/gr.088559.108
发表时间: 2009-07-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Prabhu, Snehit;Pe'er, Itsik
通讯作者: Pe'er, Itsik
DOI: 10.1006/geno.2001.6547
发表时间: 2001-06-01
期刊: GENOMICS
影响因子: 4.4
作者:
Ding, Y;Johnson, MD;Shizuya, H
通讯作者: Shizuya, H