Inflated type I error rates when using aggregation methods to analyze rare variants in the 1000 Genomes Project exon sequencing data in unrelated individuals: summary results from Group 7 at Genetic Analysis Workshop 17.

Inflated type I error rates when using aggregation methods to analyze rare variants in the 1000 Genomes Project exon sequencing data in unrelated individuals: summary results from Group 7 at Genetic Analysis Workshop 17.
复制标题

DOI:
10.1002/gepi.20650
复制
发表时间:
2011
影响因子:
2.1
通讯作者:
Pugh, Elizabeth
Pugh, Elizabeth
中科院分区:
医学4区
文献类型:
--
作者:
Tintle, Nathan;Aschard, Hugues;Hu, Inchi;Nock, Nora;Wang, Haitian;Pugh, Elizabeth

文献摘要

参考文献

被引文献

相似文献

作为遗传分析研讨会17(GAW 17)的一部分,我们的研究小组考虑了新的标准方法在下一代测序数据中分析基因型-表型关联的应用。我们的研究小组在分析GAW 17下一代测序数据时发现了一个主要问题:I型错误和假阳性报告概率高于基于经验I型错误水平的预期(高达90%)。出现两个主要原因:群体分层和罕见变异之间的长程相关性(配子相不平衡)。由于样本的多样性,预计会对人群进行分层。罕见变异之间的相关性可归因于两种随机原因(例如,25,000个标记中有近10,000个是私人变异,样本量很小[n = 697])和非随机原因(观察到的相关性比随机机会预期的更多)。主成分分析用于控制群体结构,并有助于最大限度地减少I型错误,但这是以识别更少的因果变异为代价的。一种新的多元回归方法显示出处理标记之间相关性的希望。需要进一步的工作,首先,以确定控制I型错误的测序数据分析的最佳实践,然后探索和比较许多有前途的新的聚合方法,用于识别与疾病表型相关的标记。
As part of Genetic Analysis Workshop 17 (GAW17), our group considered the application of novel and standard approaches to the analysis of genotype-phenotype association in next-generation sequencing data. Our group identified a major issue in the analysis of the GAW17 next-generation sequencing data: type I error and false-positive report probability rates higher than those expected based on empirical type I error levels (as high as 90%). Two main causes emerged: population stratification and long-range correlation (gametic phase disequilibrium) between rare variants. Population stratification was expected because of the diverse sample. Correlation between rare variants was attributable to both random causes (e.g., nearly 10,000 of 25,000 markers were private variants, and the sample size was small [n = 697]) and nonrandom causes (more correlation was observed than was expected by random chance). Principal components analysis was used to control for population structure and helped to minimize type I errors, but this was at the expense of identifying fewer causal variants. A novel multiple regression approach showed promise to handle correlation between markers. Further work is needed, first, to identify best practices for the control of type I errors in the analysis of sequencing data and then to explore and compare the many promising new aggregating approaches for identifying markers associated with disease phenotypes.
DOI: 10.1038/ng.686
发表时间: 2010-11
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --
DOI: 10.1002/gepi.20476
发表时间: 2009
影响因子: 2.1
作者:
Tintle, Nathan;Lantieri, Francesca;Lebrec, Jeremie;Sohns, Melanie;Ballard, David;Bickeboeller, Heike
通讯作者: Bickeboeller, Heike
DOI: 10.1101/gr.176601
发表时间: 2001-05-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Ng, PC;Henikoff, S
通讯作者: Henikoff, S
DOI: 10.1186/1753-6561-5-s9-s2
发表时间: 2011-11-29
期刊: BMC proceedings
影响因子: --
作者:
Almasy L;Dyer TD;Peralta JM;Kent JW Jr;Charlesworth JC;Curran JE;Blangero J
通讯作者: Blangero J
DOI: 10.1016/j.ajhg.2008.06.024
发表时间: 2008-09-12
影响因子: 9.8
作者:
Li, Bingshan;Leal, Suzanne M.
通讯作者: Leal, Suzanne M.