Quantifying the amount of missing information in genetic association studies

Quantifying the amount of missing information in genetic association studies
复制标题

DOI:
10.1002/gepi.20181
复制
发表时间:
2006-12-01
影响因子:
2.1
通讯作者:
Nicolae, Dan L.
Nicolae, Dan L.
中科院分区:
医学4区
文献类型:
--
作者:
Nicolae, Dan L.

文献摘要

被引文献

相似文献

许多遗传分析是在信息不完整的情况下进行的;例如,在基于单倍型的关联研究中相位未知。对可用信息量的度量可用于研究和/或分析的有效规划。特别是,两组标记之间的连锁不平衡(LD)可被解释为一组标记包含的用于检验第二组中 等位基因频率差异的信息量,并且测量LD可被视为对缺失数据问题中的信息进行量化。我们引入一个用于测量两组变量之间关联的框架;例如,两组不同标记的基因型数据,或者给定多态性集合的单倍型和基因型数据。目标是量化一个数据集中有多少信息,例如一组单核苷酸多态性(SNP)的基因型数据,用于估计作为第二数据集频率函数的参数,例如单倍型频率,相对于实际观察到完整数据(例如单倍型)的理想情况。在两组相互排斥的标记的基因型数据的情况下,该度量确定多位点LD的量,如果每组由一个双等位基因标记组成,则它等于经典度量r(2)。一般来说,这些度量被解释为在病例 - 对照检验中达到相同功效所需的样本量渐近比。本文的重点是病例 - 对照等位基因/单倍型检验,但该框架可轻松扩展到其他情况,如根据等位基因/单倍型计数对数量性状进行回归,或者对基因型或双倍型进行检验。我们强调该方法的应用,包括用于浏览国际人类基因组单体型图(HapMap)数据库的工具[国际人类基因组单体型图协作组,2003],以及定位克隆研究的基因分型策略。《遗传流行病学》30:703 - 717, 2006。(c) 2006威利 - 利斯公司
Many genetic analyses are done with incomplete information; for example, unknown phase in haplotype-based association studies. Measures of the amount of available information can be used for efficient planning of studies and/or analyses. In particular, the linkage disequilibrium (LD) between two sets of markers can be interpreted as the amount of information one set of markers contains for testing allele frequency differences in the second set, and measuring LD can be viewed as quantifying information in a missing data problem. We introduce a framework for measuring the association between two sets of variables; for example, genotype data for two distinct groups of markers, or haplotype and genotype data for a given set of polymorphisms. The goal is to quantify how much information is in one data set, e.g. genotype data for a set of SNPs, for estimating parameters that are functions of frequencies in the second data set, e.g. haplotype frequencies, relative to the ideal case of actually observing the complete data, e.g. haplotypes. In the case of genotype data on two mutually exclusive sets of markers, the measure determines the amount of multi-locus LD, and is equal to the classical measure r(2), if the sets consist each of one bi-allelic marker. In general, the measures are interpreted as the asymptotic ratio of sample sizes necessary to achieve the same power in case-control testing. The focus of this paper is on case-control allele/haplotype tests, but the framework can be extended easily to other settings like regressing quantitative traits on allele/haplotype counts, or tests on genotypes or diplotypes. We highlight applications of the approach, including tools for navigating the HapMap database [The International HapMap Consortium, 2003], and genotyping strategies for positional cloning studies. Genet. Epidemiol. 30:703-717, 2006. (c) 2006 Wiley-Liss, Inc.