Anchored pseudo-de novo assembly of human genomes identifies extensive sequence variation from unmapped sequence reads.

Anchored pseudo-de novo assembly of human genomes identifies extensive sequence variation from unmapped sequence reads.
复制标题

DOI:
10.1007/s00439-016-1667-5
复制
发表时间:
2016-07
期刊:
影响因子:
5.3
通讯作者:
Brown KH
Brown KH
中科院分区:
生物学2区
文献类型:
--
作者:
Faber-Hammond JJ;Brown KH

文献摘要

被引文献

相似文献

人类基因组参考(HGR)的完成标志着基因组学时代的开始,然而,尽管它的实用性,普遍的应用是有限的,在其发展中使用的个人数量少。这通过未能在HGR内映射的高质量序列读数的存在而突出。未能作图的序列通常占总读数的2- 5%,这可能包含将增强我们对群体变异、进化和疾病的理解的区域。或者,可以创建完整的从头组装,但这些有效地忽略了HGR的基础。为了找到一个中间立场,我们开发了一个生物信息学管道,将双端读段作为单独的单读段映射到HGR,导出不可映射的读段,重新组装每个个体的这些读段,然后将组装组合成用于比较分析的二级参考组装。使用45个不同的1000个基因组计划个体,我们确定了351,361个重叠群,覆盖GRCh 38中未并入的195.5 Mb序列。30,879个重叠群在多个个体中呈现,其中约40%显示出高序列复杂性。99.9%生成了基因组坐标,其中52.5%显示出高质量的作图得分。与古代人类和灵长类动物的比较基因组分析显示了显着的序列比对和与模式生物RefSeq基因数据集的比较,确定了新的人类基因。如果纳入,这些序列将扩大HGR,但更重要的是,我们的数据强调,使用这种方法,低覆盖率(约10-20×)的下一代测序仍然可以用于识别新的未映射序列,以探索有助于人类表型变异,疾病和个人基因组医学功能的生物学功能。
The human genome reference (HGR) completion marked the genomics era beginning, yet despite its utility universal application is limited by the small number of individuals used in its development. This is highlighted by the presence of high-quality sequence reads failing to map within the HGR. Sequences failing to map generally represent 2–5 % of total reads, which may harbor regions that would enhance our understanding of population variation, evolution, and disease. Alternatively, complete de novo assemblies can be created, but these effectively ignore the groundwork of the HGR. In an effort to find a middle ground, we developed a bioinformatic pipeline that maps paired-end reads to the HGR as separate single reads, exports unmappable reads, de novo assembles these reads per individual and then combines assemblies into a secondary reference assembly used for comparative analysis. Using 45 diverse 1000 Genomes Project individuals, we identified 351,361 contigs covering 195.5 Mb of sequence unincorporated in GRCh38. 30,879 contigs are represented in multiple individuals with ~40 % showing high sequence complexity. Genomic coordinates were generated for 99.9 %, with 52.5 % exhibiting high-quality mapping scores. Comparative genomic analyses with archaic humans and primates revealed significant sequence alignments and comparisons with model organism RefSeq gene datasets identified novel human genes. If incorporated, these sequences will expand the HGR, but more importantly our data highlight that with this method low coverage (~10–20×) next-generation sequencing can still be used to identify novel unmapped sequences to explore biological functions contributing to human phenotypic variation, disease and functionality for personal genomic medicine.