IMPROVING POPULATION-SPECIFIC ALLELE FREQUENCY ESTIMATES BY ADAPTING SUPPLEMENTAL DATA: AN EMPIRICAL BAYES APPROACH.

IMPROVING POPULATION-SPECIFIC ALLELE FREQUENCY ESTIMATES BY ADAPTING SUPPLEMENTAL DATA: AN EMPIRICAL BAYES APPROACH.
复制标题

DOI:
10.1214/07-aoas121
复制
发表时间:
2007-12
期刊:
The annals of applied statistics
影响因子:
--
通讯作者:
Marc Coram;Hua Tang
Marc Coram;Hua Tang
中科院分区:
其他
文献类型:
--
作者:
Marc Coram;Hua Tang

文献摘要

被引文献

相似文献

遗传标记的等位基因频率的估计是生物学和生物医学研究的关键组成部分,例如人类遗传变异或遗传性状的遗传病因学研究。随着遗传数据越来越容易获得,研究人员面临着一个困境:何时应该将其他研究和人群亚组的数据与原始数据合并?合并额外的样本通常会降低频率估计值的方差;然而,如果使用不当,由于人群分层,合并的估计值可能会严重偏倚。由于这种潜在的偏见,大多数研究人员避免合并,即使是具有相同种族背景和居住在同一大陆的样本。在这里,我们提出了一个经验贝叶斯方法估计等位基因频率的单核苷酸多态性。该过程自适应地结合了相关样本的基因型,使得更相似的样本对估计值有更大的影响。在我们考虑的每个例子中,我们的估计器都实现了比池化或不池化都小的均方误差(MSE),有时在两个极端情况下都有很大的改善。正如与真实的数据示例仔细匹配的模拟研究所示,引入的偏差很小。我们的方法是特别有用的,当小群体的个人基因型在大量的标记,我们很可能会遇到在全基因组关联研究的情况。
Estimation of the allele frequency at genetic markers is a key ingredient in biological and biomedical research, such as studies of human genetic variation or of the genetic etiology of heritable traits. As genetic data becomes increasingly available, investigators face a dilemma: when should data from other studies and population subgroups be pooled with the primary data? Pooling additional samples will generally reduce the variance of the frequency estimates; however, used inappropriately, pooled estimates can be severely biased due to population stratification. Because of this potential bias, most investigators avoid pooling, even for samples with the same ethnic background and residing on the same continent. Here, we propose an empirical Bayes approach for estimating allele frequencies of single nucleotide polymorphisms. This procedure adaptively incorporates genotypes from related samples, so that more similar samples have a greater influence on the estimates. In every example we have considered, our estimator achieves a mean squared error (MSE) that is smaller than either pooling or not, and sometimes substantially improves over both extremes. The bias introduced is small, as is shown by a simulation study that is carefully matched to a real data example. Our method is particularly useful when small groups of individuals are genotyped at a large number of markers, a situation we are likely to encounter in a genome-wide association study.