Relationship estimation from whole-genome sequence data.

Relationship estimation from whole-genome sequence data.
复制标题

DOI:
10.1371/journal.pgen.1004144
复制
发表时间:
2014-01
期刊:
影响因子:
4.5
通讯作者:
Huff CD
Huff CD
中科院分区:
生物学2区
文献类型:
--
作者:
Li H;Glusman G;Hu H;Shankaracharya;Caballero J;Hubley R;Witherspoon D;Guthery SL;Mauldin DE;Jorde LB;Hood L;Roach JC;Huff CD

文献摘要

参考文献

被引文献

相似文献

确定一对个体之间的关系是遗传学的基本应用。此前,我们和其他人已经证明,从高密度单核苷酸多态性(SNP)数据生成的血统身份(IBD)信息可以极大地提高遗传关系检测的能力和准确性。全基因组测序 (WGS) 标志着通过分析所有单核苷酸变异 (SNV) 来增加遗传标记密度的最后一步,因此有可能通过更准确地检测 IBD 片段和更精确地解析 IBD 片段边界来进一步改进关系检测。然而,WGS 引入了新的复杂性,为了实现关系检测的这些改进,必须解决这些复杂性。为了评估这些复杂性,我们根据 WGS 数据估计了 30 个家族 258 个个体中 1490 个已知配对关系的遗传关系,以及作为对照的 46 个群体样本。我们使用三种已建立的 IBD 方法:GERMLINE、fastIBD 和 ISCA,在谱系和对照数据集中鉴定了几个具有过量成对 IBD 的基因组区域。与高密度微阵列数据集相比,这些虚假 IBD 片段使对照之间检测到的假阳性关系的比率增加了 10 倍。为了解决这个问题,我们开发了一种新方法来识别和掩盖 IBD 过多的基因组区域。该方法在ERSA 2.0中实现,完全解决了神秘关系检测率过高的问题,同时提高了关系估计的准确性。 ERSA 2.0 检测了 30 个家庭中的所有 1 级至 6 级关系,以及 55% 的 9 级至 11 级关系。我们估计,相对于远距离关系的高密度微阵列数据,WGS 数据的关系检测能力提高了 5% 至 15%。我们的结果确定了 IBD 绘图中存在很大问题的基因组区域,并引入了新软件,可以从全基因组序列数据中准确检测 1 级到 9 级关系。确定一对个体之间的关系是遗传学的基本应用。最准确的关系估计方法依赖于对个体之间遗传共享的精确、局部估计。早期的方法是根据高密度遗传标记数据生成这些估计值。我们使用全基因组序列数据对 30 个家族的 258 个个体之间的 1490 个已知配对关系以及作为对照的 46 个群体样本进行了关系估计。我们的结果表明,全基因组测序特有的复杂性导致基因组区域容易出现遗传共享的假阳性估计。我们提供了这些虚假 IBD 区域的地图,并引入了在软件包 ERSA 2.0 中实施的新方法来控制虚假 IBD。我们表明,相对于高密度遗传标记数据,ERSA 2.0 对全基因组序列数据的远距离关系的关系检测能力提高了 5% 至 15%。
The determination of the relationship between a pair of individuals is a fundamental application of genetics. Previously, we and others have demonstrated that identity-by-descent (IBD) information generated from high-density single-nucleotide polymorphism (SNP) data can greatly improve the power and accuracy of genetic relationship detection. Whole-genome sequencing (WGS) marks the final step in increasing genetic marker density by assaying all single-nucleotide variants (SNVs), and thus has the potential to further improve relationship detection by enabling more accurate detection of IBD segments and more precise resolution of IBD segment boundaries. However, WGS introduces new complexities that must be addressed in order to achieve these improvements in relationship detection. To evaluate these complexities, we estimated genetic relationships from WGS data for 1490 known pairwise relationships among 258 individuals in 30 families along with 46 population samples as controls. We identified several genomic regions with excess pairwise IBD in both the pedigree and control datasets using three established IBD methods: GERMLINE, fastIBD, and ISCA. These spurious IBD segments produced a 10-fold increase in the rate of detected false-positive relationships among controls compared to high-density microarray datasets. To address this issue, we developed a new method to identify and mask genomic regions with excess IBD. This method, implemented in ERSA 2.0, fully resolved the inflated cryptic relationship detection rates while improving relationship estimation accuracy. ERSA 2.0 detected all 1st through 6th degree relationships, and 55% of 9th through 11th degree relationships in the 30 families. We estimate that WGS data provides a 5% to 15% increase in relationship detection power relative to high-density microarray data for distant relationships. Our results identify regions of the genome that are highly problematic for IBD mapping and introduce new software to accurately detect 1st through 9th degree relationships from whole-genome sequence data. The determination of the relationship between a pair of individuals is a fundamental application of genetics. The most accurate methods for relationship estimation rely on precise, localized estimates of genetic sharing between individuals. Earlier methods have generated these estimates from high-density genetic marker data. We performed relationship estimation using whole-genome sequence data for 1490 known pairwise relationships among 258 individuals in 30 families along with 46 population samples as controls. Our results demonstrate that complexities specific to whole-genome sequencing result in regions of the genome that are prone to false-positive estimates of genetic sharing. We provide a map of these spurious IBD regions and introduce new methods, implemented in the software package ERSA 2.0, to control for spurious IBD. We show that ERSA 2.0 provides a 5% to 15% increase in relationship detection power for distant relationships with whole-genome sequence data relative to high-density genetic marker data.
DOI: 10.1038/nature06258
发表时间: 2007-10-18
期刊: NATURE
影响因子: 64.8
作者:
Frazer, Kelly A.;Ballinger, Dennis G.;Cox, David R.;Hinds, David A.;Stuve, Laura L.;Gibbs, Richard A.;Belmont, John W.;Boudreau, Andrew;Hardenbol, Paul;Leal, Suzanne M.;Pasternak, Shiran;Wheeler, David A.;Willis, Thomas D.;Yu, Fuli;Yang, Huanming;Zeng, Changqing;Gao, Yang;Hu, Haoran;Hu, Weitao;Li, Chaohua;Lin, Wei;Liu, Siqi;Pan, Hao;Tang, Xiaoli;Wang, Jian;Wang, Wei;Yu, Jun;Zhang, Bo;Zhang, Qingrun;Zhao, Hongbin;Zhao, Hui;Zhou, Jun;Gabriel, Stacey B.;Barry, Rachel;Blumenstiel, Brendan;Camargo, Amy;Defelice, Matthew;Faggart, Maura;Goyette, Mary;Gupta, Supriya;Moore, Jamie;Nguyen, Huy;Onofrio, Robert C.;Parkin, Melissa;Roy, Jessica;Stahl, Erich;Winchester, Ellen;Ziaugra, Liuda;Altshuler, David;Shen, Yan;Yao, Zhijian;Huang, Wei;Chu, Xun;He, Yungang;Jin, Li;Liu, Yangfan;Shen, Yayun;Sun, Weiwei;Wang, Haifeng;Wang, Yi;Wang, Ying;Xiong, Xiaoyan;Xu, Liang;Waye, Mary M. Y.;Tsui, Stephen K. W.;Wong, J. Tze-Fei;Galver, Luana M.;Fan, Jian-Bing;Gunderson, Kevin;Murray, Sarah S.;Oliphant, Arnold R.;Chee, Mark S.;Montpetit, Alexandre;Chagnon, Fanny;Ferretti, Vincent;Leboeuf, Martin;Olivier, Jean-Franccois;Phillips, Michael S.;Roumy, Stephanie;Sallee, Clementine;Verner, Andrei;Hudson, Thomas J.;Kwok, Pui-Yan;Cai, Dongmei;Koboldt, Daniel C.;Miller, Raymond D.;Pawlikowska, Ludmila;Taillon-Miller, Patricia;Xiao, Ming;Tsui, Lap-Chee;Mak, William;Song, You Qiang;Tam, Paul K. H.;Nakamura, Yusuke;Kawaguchi, Takahisa;Kitamoto, Takuya;Morizono, Takashi;Nagashima, Atsushi;Ohnishi, Yozo;Sekine, Akihiro;Tanaka, Toshihiro;Tsunoda, Tatsuhiko;Deloukas, Panos;Bird, Christine P.;Delgado, Marcos;Dermitzakis, Emmanouil T.;Gwilliam, Rhian;Hunt, Sarah;Morrison, Jonathan;Powell, Don;Stranger, Barbara E.;Whittaker, Pamela;Bentley, David R.;Daly, Mark J.;de Bakker, Paul I. W.;Barrett, Jeff;Chretien, Yves R.;Maller, Julian;McCarroll, Steve;Patterson, Nick;Pe'er, Itsik;Price, Alkes;Purcell, Shaun;Richter, Daniel J.;Sabeti, Pardis;Saxena, Richa;Schaffner, Stephen F.;Sham, Pak C.;Varilly, Patrick;Altshuler, David;Stein, Lincoln D.;Krishnan, Lalitha;Smith, Albert Vernon;Tello-Ruiz, Marcela K.;Thorisson, Gudmundur A.;Chakravarti, Aravinda;Chen, Peter E.;Cutler, David J.;Kashuk, Carl S.;Lin, Shin;Abecasis, Goncalo R.;Guan, Weihua;Li, Yun;Munro, Heather M.;Qin, Zhaohui Steve;Thomas, Daryl J.;McVean, Gilean;Auton, Adam;Bottolo, Leonardo;Cardin, Niall;Eyheramendy, Susana;Freeman, Colin;Marchini, Jonathan;Myers, Simon;Spencer, Chris;Stephens, Matthew;Donnelly, Peter;Cardon, Lon R.;Clarke, Geraldine;Evans, David M.;Morris, Andrew P.;Weir, Bruce S.;Tsunoda, Tatsuhiko;Johnson, Todd A.;Mullikin, James C.;Sherry, Stephen T.;Feolo, Michael;Skol, Andrew
通讯作者: Skol, Andrew
DOI: 10.1126/science.1092500
发表时间: 2004-04-23
期刊: SCIENCE
影响因子: 56.9
作者:
McVean, GAT;Myers, SR;Donnelly, P
通讯作者: Donnelly, P
DOI: 10.1101/gr.081398.108
发表时间: 2009-02-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Gusev, Alexander;Lowe, Jennifer K.;Pe'er, Itsik
通讯作者: Pe'er, Itsik
DOI: 10.1111/j.1469-1809.1975.tb00120.x
发表时间: 1975-01-01
影响因子: 1.9
作者:
THOMPSON, EA
通讯作者: THOMPSON, EA
DOI: 10.1534/genetics.113.150029
发表时间: 2013-06
期刊: Genetics
影响因子: 3.3
作者:
Browning BL;Browning SR
通讯作者: Browning SR