Construction of JRG (Japanese reference genome) with single-molecule real-time sequencing

Construction of JRG (Japanese reference genome) with single-molecule real-time sequencing
复制标题

DOI:
10.1038/s41439-019-0057-7
复制
发表时间:
2019-06
影响因子:
1.5
通讯作者:
Masao Nagasaki;Y. Kuroki;Tomoko F. Shibata;F. Katsuoka;Takahiro Mimori;Y. Kawai;N. Minegishi;A. Hozawa;S. Kuriyama;Yoichi Suzuki;H. Kawame;Fuji Nagami;Takako Takai-Igarashi;S. Ogishima;Kaname Kojima;K. Misawa;Osamu Tanabe;N. Fuse;Hiroshi Tanaka;N. Yaegashi;K. Kinoshita;Shiego Kure;J. Yasuda;Masayuki Yamamoto
Masao Nagasaki;Y. Kuroki;Tomoko F. Shibata;F. Katsuoka;Takahiro Mimori;Y. Kawai;N. Minegishi;A. Hozawa;S. Kuriyama;Yoichi Suzuki;H. Kawame;Fuji Nagami;Takako Takai-Igarashi;S. Ogishima;Kaname Kojima;K. Misawa;Osamu Tanabe;N. Fuse;Hiroshi Tanaka;N. Yaegashi;K. Kinoshita;Shiego Kure;J. Yasuda;Masayuki Yamamoto
中科院分区:
--
文献类型:
--
作者:
Masao Nagasaki;Y. Kuroki;Tomoko F. Shibata;F. Katsuoka;Takahiro Mimori;Y. Kawai;N. Minegishi;A. Hozawa;S. Kuriyama;Yoichi Suzuki;H. Kawame;Fuji Nagami;Takako Takai-Igarashi;S. Ogishima;Kaname Kojima;K. Misawa;Osamu Tanabe;N. Fuse;Hiroshi Tanaka;N. Yaegashi;K. Kinoshita;Shiego Kure;J. Yasuda;Masayuki Yamamoto

文献摘要

相似文献

在最近的基因组分析中,特定人群的参考面板表明很重要。然而,基于短读长测序数据的参考面板不足以覆盖长插入。因此,长插入的性质尚未得到充分记录。在这里,我们使用单分子实时测序数据组装了日本基因组,并对组装基因组中发现的插入进行了表征。相对于国际参考序列 (GRCh38),我们在组装的基因组中鉴定了 3691 个插入,范围从 100bps 到约 10,000bps。为了验证和表征这些插入,我们将来自 1070 个日本个体和来自其他 8 个群体的 728 个个体的短读映射到整合到 GRCh38 中的插入。有了这个结果,我们通过将至少两个日本人共享的 903 个经过验证的插入(总计 1,086,173 个碱基)整合到 GRCh38 中,构建了 JRGv1(日本参考基因组版本 1)。我们还通过连接 3559 个经过验证的插入片段(总共 2,536,870 个碱基)构建了诱饵JRGv1,这些插入片段由至少两个日本人或六个其他程序集共享。该组件平均将对准率提高了 0.4%。这些结果证明了完善参考组装和创建特定人群参考基因组的重要性。 JRGv1 和 decoyJRGv1 可在 JRG 网站上获取。
In recent genome analyses, population-specific reference panels have indicated important. However, reference panels based on short-read sequencing data do not sufficiently cover long insertions. Therefore, the nature of long insertions has not been well documented. Here, we assembled a Japanese genome using single-molecule real-time sequencing data and characterized insertions found in the assembled genome. We identified 3691 insertions ranging from 100 bps to ~10,000 bps in the assembled genome relative to the international reference sequence (GRCh38). To validate and characterize these insertions, we mapped short-reads from 1070 Japanese individuals and 728 individuals from eight other populations to insertions integrated into GRCh38. With this result, we constructed JRGv1 (Japanese Reference Genome version 1) by integrating the 903 verified insertions, totaling 1,086,173 bases, shared by at least two Japanese individuals into GRCh38. We also constructed decoyJRGv1 by concatenating 3559 verified insertions, totaling 2,536,870 bases, shared by at least two Japanese individuals or by six other assemblies. This assembly improved the alignment ratio by 0.4% on average. These results demonstrate the importance of refining the reference assembly and creating a population-specific reference genome. JRGv1 and decoyJRGv1 are available at the JRG website.