Extensive sequencing of seven human genomes to characterize benchmark reference materials.

Extensive sequencing of seven human genomes to characterize benchmark reference materials.
复制标题

DOI:
10.1038/sdata.2016.25
复制
发表时间:
2016-06-07
期刊:
影响因子:
9.8
通讯作者:
Salit M
Salit M
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Zook JM;Catoe D;McDaniel J;Vang L;Spies N;Sidow A;Weng Z;Liu Y;Mason CE;Alexander N;Henaff E;McIntyre AB;Chandramohan D;Chen F;Jaeger E;Moshrefi A;Pham K;Stedman W;Liang T;Saghbini M;Dzakula Z;Hastie A;Cao H;Deikus G;Schadt E;Sebra R;Bashir A;Truty RM;Chang CC;Gulbahce N;Zhao K;Ghosh S;Hyland F;Fu Y;Chaisson M;Xiao C;Trow J;Sherry ST;Zaranek AW;Ball M;Bobe J;Estep P;Church GM;Marks P;Kyriazopoulou-Panagiotopoulou S;Zheng GX;Schnall-Levin M;Ordonez HS;Mudivarti PA;Giorda K;Sheng Y;Rypdal KB;Salit M

文献摘要

被引文献

相似文献

由美国国家标准与技术研究所(NIST)主办的瓶中基因组联盟正在为人类基因组测序创建参考材料和数据,以及基因组比较和基准测试方法。在这里,我们描述了7个人类基因组的一个大型,多样化的测序数据集;其中5个是当前或候选的NIST参考材料。先导基因组NA 12878已作为NIST RM 8398发布。我们还描述了来自两个个人基因组计划三人组的数据,一个是德系犹太血统,一个是中国血统。这些数据来自12种技术:BioNano Genomics,Complete Genomics配对末端和LFR,Ion Proton外显子组,Oxford Nanopore,Pacific Biosciences,SOLiD,10 X Genomics GemCode WGS,以及Illumina外显子组和WGS配对末端,配对和合成长读段。这些人的细胞系、DNA和数据是公开的。因此,我们希望这些数据有助于揭示有关人类基因组的新信息,并改善测序技术,SNP,indel和结构变体调用以及从头组装。
The Genome in a Bottle Consortium, hosted by the National Institute of Standards and Technology (NIST) is creating reference materials and data for human genome sequencing, as well as methods for genome comparison and benchmarking. Here, we describe a large, diverse set of sequencing data for seven human genomes; five are current or candidate NIST Reference Materials. The pilot genome, NA12878, has been released as NIST RM 8398. We also describe data from two Personal Genome Project trios, one of Ashkenazim Jewish ancestry and one of Chinese ancestry. The data come from 12 technologies: BioNano Genomics, Complete Genomics paired-end and LFR, Ion Proton exome, Oxford Nanopore, Pacific Biosciences, SOLiD, 10X Genomics GemCode WGS, and Illumina exome and WGS paired-end, mate-pair, and synthetic long reads. Cell lines, DNA, and data from these individuals are publicly available. Therefore, we expect these data to be useful for revealing novel information about the human genome and improving sequencing technologies, SNP, indel, and structural variant calling, and de novo assembly.