Evaluating the effective numbers of independent tests and significant p-value thresholds in commercial genotyping arrays and public imputation reference datasets.

Evaluating the effective numbers of independent tests and significant p-value thresholds in commercial genotyping arrays and public imputation reference datasets.
复制标题

DOI:
10.1007/s00439-011-1118-2
复制
发表时间:
2012-05
期刊:
影响因子:
5.3
通讯作者:
Sham PC
Sham PC
中科院分区:
生物学2区
文献类型:
--
作者:
Li MX;Yeung JM;Cherny SS;Sham PC

文献摘要

参考文献

被引文献

相似文献

目前的全基因组关联研究(GWAS)使用商业基因分型微阵列,可以分析超过一百万个单核苷酸多态性(SNP)。先进的统计基因型插补算法和参考人群的大型SNP数据库进一步增加了SNP的数量。在这种全基因组研究中,在解释统计学显著性时需要考虑大量SNP的测试,但由于连锁不平衡(LD),SNP的非独立性使其复杂化。以前的几个小组已经提出了使用的有效数量的独立标记(M e)的多重测试的调整,但目前的方法计算M e是有限的精度或计算速度。在这里,我们报告一个更强大和快速的方法来计算M e。应用这种有效的方法[在名为Genetic type 1 error calculator(GEC)的免费软件工具中实现],我们系统地检查了13个Illumina或Affyssin基因分型阵列的M e和将全基因组1型错误率控制在0.05所需的相应p值阈值,以及HapMap项目和1000个基因组项目数据集,这些数据集广泛用于基因型插补作为参考组。我们的研究结果表明,对于早期商业化基因分型阵列,使用约10−7的p值阈值作为全基因组显著性的标准,但对于当前或合并的商业化基因分型阵列,p值阈值稍微更严格一些,约为5 × 10−8,对于1000个基因组计划数据集中的所有常见SNP,约为10−8,对于仅在基因内的常见SNP,约为5 × 10−8。本文的在线版本(doi:10.1007/s 00439 -011-1118-2)包含补充材料,可供授权用户使用。
Current genome-wide association studies (GWAS) use commercial genotyping microarrays that can assay over a million single nucleotide polymorphisms (SNPs). The number of SNPs is further boosted by advanced statistical genotype-imputation algorithms and large SNP databases for reference human populations. The testing of a huge number of SNPs needs to be taken into account in the interpretation of statistical significance in such genome-wide studies, but this is complicated by the non-independence of SNPs because of linkage disequilibrium (LD). Several previous groups have proposed the use of the effective number of independent markers (M e) for the adjustment of multiple testing, but current methods of calculation for M e are limited in accuracy or computational speed. Here, we report a more robust and fast method to calculate M e. Applying this efficient method [implemented in a free software tool named Genetic type 1 error calculator (GEC)], we systematically examined the M e, and the corresponding p-value thresholds required to control the genome-wide type 1 error rate at 0.05, for 13 Illumina or Affymetrix genotyping arrays, as well as for HapMap Project and 1000 Genomes Project datasets which are widely used in genotype imputation as reference panels. Our results suggested the use of a p-value threshold of ~10−7 as the criterion for genome-wide significance for early commercial genotyping arrays, but slightly more stringent p-value thresholds ~5 × 10−8 for current or merged commercial genotyping arrays, ~10−8 for all common SNPs in the 1000 Genomes Project dataset and ~5 × 10−8 for the common SNPs only within genes. The online version of this article (doi:10.1007/s00439-011-1118-2) contains supplementary material, which is available to authorized users.
DOI: 10.1038/nature06258
发表时间: 2007-10-18
期刊: NATURE
影响因子: 64.8
作者:
Frazer, Kelly A.;Ballinger, Dennis G.;Cox, David R.;Hinds, David A.;Stuve, Laura L.;Gibbs, Richard A.;Belmont, John W.;Boudreau, Andrew;Hardenbol, Paul;Leal, Suzanne M.;Pasternak, Shiran;Wheeler, David A.;Willis, Thomas D.;Yu, Fuli;Yang, Huanming;Zeng, Changqing;Gao, Yang;Hu, Haoran;Hu, Weitao;Li, Chaohua;Lin, Wei;Liu, Siqi;Pan, Hao;Tang, Xiaoli;Wang, Jian;Wang, Wei;Yu, Jun;Zhang, Bo;Zhang, Qingrun;Zhao, Hongbin;Zhao, Hui;Zhou, Jun;Gabriel, Stacey B.;Barry, Rachel;Blumenstiel, Brendan;Camargo, Amy;Defelice, Matthew;Faggart, Maura;Goyette, Mary;Gupta, Supriya;Moore, Jamie;Nguyen, Huy;Onofrio, Robert C.;Parkin, Melissa;Roy, Jessica;Stahl, Erich;Winchester, Ellen;Ziaugra, Liuda;Altshuler, David;Shen, Yan;Yao, Zhijian;Huang, Wei;Chu, Xun;He, Yungang;Jin, Li;Liu, Yangfan;Shen, Yayun;Sun, Weiwei;Wang, Haifeng;Wang, Yi;Wang, Ying;Xiong, Xiaoyan;Xu, Liang;Waye, Mary M. Y.;Tsui, Stephen K. W.;Wong, J. Tze-Fei;Galver, Luana M.;Fan, Jian-Bing;Gunderson, Kevin;Murray, Sarah S.;Oliphant, Arnold R.;Chee, Mark S.;Montpetit, Alexandre;Chagnon, Fanny;Ferretti, Vincent;Leboeuf, Martin;Olivier, Jean-Franccois;Phillips, Michael S.;Roumy, Stephanie;Sallee, Clementine;Verner, Andrei;Hudson, Thomas J.;Kwok, Pui-Yan;Cai, Dongmei;Koboldt, Daniel C.;Miller, Raymond D.;Pawlikowska, Ludmila;Taillon-Miller, Patricia;Xiao, Ming;Tsui, Lap-Chee;Mak, William;Song, You Qiang;Tam, Paul K. H.;Nakamura, Yusuke;Kawaguchi, Takahisa;Kitamoto, Takuya;Morizono, Takashi;Nagashima, Atsushi;Ohnishi, Yozo;Sekine, Akihiro;Tanaka, Toshihiro;Tsunoda, Tatsuhiko;Deloukas, Panos;Bird, Christine P.;Delgado, Marcos;Dermitzakis, Emmanouil T.;Gwilliam, Rhian;Hunt, Sarah;Morrison, Jonathan;Powell, Don;Stranger, Barbara E.;Whittaker, Pamela;Bentley, David R.;Daly, Mark J.;de Bakker, Paul I. W.;Barrett, Jeff;Chretien, Yves R.;Maller, Julian;McCarroll, Steve;Patterson, Nick;Pe'er, Itsik;Price, Alkes;Purcell, Shaun;Richter, Daniel J.;Sabeti, Pardis;Saxena, Richa;Schaffner, Stephen F.;Sham, Pak C.;Varilly, Patrick;Altshuler, David;Stein, Lincoln D.;Krishnan, Lalitha;Smith, Albert Vernon;Tello-Ruiz, Marcela K.;Thorisson, Gudmundur A.;Chakravarti, Aravinda;Chen, Peter E.;Cutler, David J.;Kashuk, Carl S.;Lin, Shin;Abecasis, Goncalo R.;Guan, Weihua;Li, Yun;Munro, Heather M.;Qin, Zhaohui Steve;Thomas, Daryl J.;McVean, Gilean;Auton, Adam;Bottolo, Leonardo;Cardin, Niall;Eyheramendy, Susana;Freeman, Colin;Marchini, Jonathan;Myers, Simon;Spencer, Chris;Stephens, Matthew;Donnelly, Peter;Cardon, Lon R.;Clarke, Geraldine;Evans, David M.;Morris, Andrew P.;Weir, Bruce S.;Tsunoda, Tatsuhiko;Johnson, Todd A.;Mullikin, James C.;Sherry, Stephen T.;Feolo, Michael;Skol, Andrew
通讯作者: Skol, Andrew
DOI: 10.1007/bf01245622
发表时间: 1968-01-01
影响因子: 5.4
作者:
HILL W G;ROBERTSON A
通讯作者: ROBERTSON A
DOI: 10.1002/gepi.20310
发表时间: 2008-05-01
影响因子: 2.1
作者:
Gao, Xiaoyi;Stamier, Joshua;Martin, Eden R.
通讯作者: Martin, Eden R.
DOI: 10.1093/bioinformatics/bti689
发表时间: 2005-12-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Montana, G
通讯作者: Montana, G
DOI: 10.1038/sj.hdy.6800717
发表时间: 2005-09-01
期刊: HEREDITY
影响因子: 3.8
作者:
Li, J;Ji, L
通讯作者: Ji, L