Genomic control for association studies: a semiparametric test to detect excess-haplotype sharing.

Genomic control for association studies: a semiparametric test to detect excess-haplotype sharing.
复制标题

DOI:
10.1093/biostatistics/1.4.369
复制
发表时间:
2000-12-01
期刊:
Biostatistics (Oxford, England)
影响因子:
--
通讯作者:
Wasserman, L
Wasserman, L
中科院分区:
其他
文献类型:
--
作者:
Devlin, B;Roeder, K;Wasserman, L

文献摘要

被引文献

相似文献

从共同祖先共享疾病突变的个体通常在邻近突变的遗传标记处共享等位基因,即使共同祖先是遥远的。这些相邻标记上的等位基因,称为单倍型,可以被看作是一串随机变量的实现,当个体以某种方式相关时,这些随机变量可能是依赖的。理想情况下,对于一个样本的个人都有相同的(遗传)疾病,这种依赖性-测量单倍型共享-将在疾病基因附近比在基因组的其他区域更大。在本文中,我们提出了一个半参数检验单倍型共享。我们开始通过开发一个模型,假设祖先单倍型是已知的,因此可以明确地确定来自共同祖先的单倍型共享的程度。远离疾病的标记物处的重叠量被视为具有未知分布F的随机变量,我们非参数地估计该分布F。疾病基因周围标记物的重叠被建模为混合物pF(x -θ)+(1 - p)F(x),其中p是具有疾病突变的受试者的分数。检测一个疾病基因就等于检测p是否= 0。接下来,我们放弃祖先单倍型已知的假设。为了检测单倍型的过度聚类,我们测量一组单倍型的成对重叠。在更简单的场景中,该分布被建模为位置偏移混合。为了检验假设,我们构造了一个简单的极限分布的分数检验。
Individuals who share a disease mutation from a common ancestor often share alleles at genetic markers adjacent to the mutation, even if the common ancestor is remote. The alleles at these adjacent markers, called the haplotype, can be visualized as a string of realizations of random variables, which may be dependent when individuals are related in some fashion. Ideally, for a sample of individuals all having the same (genetic) disease, this dependence-measured as haplotype-sharing-will be greater in the vicinity of disease genes than in other regions of the genome. In this paper we present a semiparametric test for haplotype-sharing. We begin by developing a model assuming that the ancestral haplotype is known and thus the extent of haplotype-sharing from a common ancestor can be determined unambiguously. The amount of overlap at markers far from the disease is treated as a random variable with an unknown distribution F, which we estimate non-parametrically. Overlap of markers surrounding disease genes are modeled as a mixture pF(x - theta) + (1 - p)F(x), in which p is the fraction of subjects with the disease mutation. Testing for a disease gene then amounts to testing whether p = 0. Next we drop the assumption that the ancestral haplotype is known. To detect excess clustering of haplotypes, we measure the pairwise overlap of a set of haplotypes. As in the simpler scenario, this distribution is modeled as a location-shift mixture. To test the hypothesis we construct a score test with a simple limiting distribution.