Direct determination of diploid genome sequences

Direct determination of diploid genome sequences
复制标题

DOI:
10.1101/070425
复制
发表时间:
2016-08
期刊:
影响因子:
7
通讯作者:
N. Weisenfeld;Vijay Kumar;Preyas Shah;D. Church;D. Jaffe
N. Weisenfeld;Vijay Kumar;Preyas Shah;D. Church;D. Jaffe
中科院分区:
生物学1区
文献类型:
--
作者:
N. Weisenfeld;Vijay Kumar;Preyas Shah;D. Church;D. Jaffe

文献摘要

被引文献

相似文献

确定生物体的基因组序列具有挑战性,但对于理解其生物学至关重要。在过去的十年中,成千上万的人类基因组被测序,为生物医学研究做出了巨大贡献。在绝大多数情况下,这些已经通过将序列读数与单个参考基因组比对来分析,从而使所得分析产生偏差,并且通常未能捕获给定基因组的新序列。已经构建了一些从头组装,没有参考偏倚,但几乎所有的组装都是通过将同源基因座合并到单个“共有”序列中来构建的,通常在自然界中不存在。这些组合不能正确地代表个体的二倍体生物学。在两种情况下,真正的二倍体从头组装已经取得了巨大的代价。一个是使用桑格测序生成的,另一个是使用数千个克隆池生成的。在这里,我们展示了一个简单的和低成本的方法来创建真正的二倍体从头组装。我们使用10 x Genomics微流控平台从~1 ng高分子量DNA制备单个文库以分配基因组。我们将这项技术应用于七个人类样本,生成低成本的HiSeq X数据,然后使用一种新的“搜索”算法Supernova组装这些数据。每次计算在一台服务器上需要两天时间。每一个都产生了长于100 kb的重叠群,长于2.5 Mb的相位块和长于15 Mb的支架。我们的方法提供了一种可扩展的能力,用于确定样品中的实际二倍体基因组序列,为基因组生物学和医学的新方法打开了大门。
Determining the genome sequence of an organism is challenging, yet fundamental to understanding its biology. Over the past decade, thousands of human genomes have been sequenced, contributing deeply to biomedical research. In the vast majority of cases, these have been analyzed by aligning sequence reads to a single reference genome, biasing the resulting analyses and, in general, failing to capture sequences novel to a given genome. Some de novo assemblies have been constructed, free of reference bias, but nearly all were constructed by merging homologous loci into single ‘consensus’ sequences, generally absent from nature. These assemblies do not correctly represent the diploid biology of an individual. In exactly two cases, true diploid de novo assemblies have been made, at great expense. One was generated using Sanger sequencing and one using thousands of clone pools. Here we demonstrate a straightforward and low-cost method for creating true diploid de novo assemblies. We make a single library from ~1 ng of high molecular weight DNA, using the 10x Genomics microfluidic platform to partition the genome. We applied this technique to seven human samples, generating low-cost HiSeq X data, then assembled these using a new ‘pushbutton’ algorithm, Supernova. Each computation took two days on a single server. Each yielded contigs longer than 100 kb, phase blocks longer than 2.5 Mb, and scaffolds longer than 15 Mb. Our method provides a scalable capability for determining the actual diploid genome sequence in a sample, opening the door to new approaches in genomic biology and medicine.