Read clouds uncover variation in complex regions of the human genome.

Read clouds uncover variation in complex regions of the human genome.
复制标题

DOI:
10.1101/gr.191189.115
复制
发表时间:
2015-10
期刊:
影响因子:
7
通讯作者:
Batzoglou S
Batzoglou S
中科院分区:
生物学1区
文献类型:
--
作者:
Bishara A;Liu Y;Weng Z;Kashef-Haghighi D;Newburger DE;West R;Sidow A;Batzoglou S

文献摘要

被引文献

相似文献

尽管越来越多的人类基因变异正在被识别和记录,但确定人类基因组重复序列中的变异仍然是一个挑战。因此,大多数群体和全基因组关联研究无法考虑这些区域的变异。问题的核心是缺乏一种测序技术来产生足够长和准确的读数,从而实现独特的作图。在这里,我们提出了一种新的方法,使用读取云,通过对来自长片段文库的DNA进行准确的短读测序,自信地对重复区域内的短读取进行比对,并实现准确的变异发现。我们的新算法,随机场对齐(RFA),通过马尔可夫随机场来捕捉受长读过程支配的短读之间的关系。我们使用了Illumina TruSeq合成长读协议的一个修改版本,该协议产生了浅序列读云。我们通过广泛的模拟测试RFA,并将其应用于发现NA12878人类样本上的变异,对于NA12878人样本,可以获得浅TruSeq Read Cloud测序数据,以及我们使用相同方法测序的浸润性乳腺癌基因组。我们证明,RFA有助于准确恢复人类基因组155Mb的变异,包括67Mb片段重复序列中94%的变异和11Mb转录序列中96%的变异,这些序列目前对短读写技术是隐藏的。
Although an increasing amount of human genetic variation is being identified and recorded, determining variants within repeated sequences of the human genome remains a challenge. Most population and genome-wide association studies have therefore been unable to consider variation in these regions. Core to the problem is the lack of a sequencing technology that produces reads with sufficient length and accuracy to enable unique mapping. Here, we present a novel methodology of using read clouds, obtained by accurate short-read sequencing of DNA derived from long fragment libraries, to confidently align short reads within repeat regions and enable accurate variant discovery. Our novel algorithm, Random Field Aligner (RFA), captures the relationships among the short reads governed by the long read process via a Markov Random Field. We utilized a modified version of the Illumina TruSeq synthetic long-read protocol, which yielded shallow-sequenced read clouds. We test RFA through extensive simulations and apply it to discover variants on the NA12878 human sample, for which shallow TruSeq read cloud sequencing data are available, and on an invasive breast carcinoma genome that we sequenced using the same method. We demonstrate that RFA facilitates accurate recovery of variation in 155 Mb of the human genome, including 94% of 67 Mb of segmental duplication sequence and 96% of 11 Mb of transcribed sequence, that are currently hidden from short-read technologies.