Genome association studies of complex diseases by case-control designs

Genome association studies of complex diseases by case-control designs
复制标题

DOI:
10.1086/373966
复制
发表时间:
2003-04-01
影响因子:
9.8
通讯作者:
Knapp, M
Knapp, M
中科院分区:
生物学1区
文献类型:
--
作者:
Fan, RZ;Knapp, M

文献摘要

被引文献

相似文献

对遗传性状进行连锁不平衡(LD)定位的一种方法是使用单标记。由于密集的标记图谱--如单核苷酸多态和高分辨率微卫星图谱--已经存在,因此将单标记LD图谱推广到高分辨率单倍型或多标记LD图谱是自然和实用的。本文研究了基于单倍型图谱或微卫星标记图谱的复杂疾病的高分辨率LD作图方法。其目的是探索结合来自单倍型块或多标记的信息的测试统计。基于两种编码方法,即基因型编码和单倍型编码,提出了Hotling的T-2统计量T-G和T-H来检验一个疾病基因座与两个单倍型区块或两个标记之间的关联。理论计算证明了这两个T-2统计量的有效性。通过简单地将两个单倍型区块的X(2)检验统计量加在一起,引入了统计量T-C,它是传统单倍型频率比较方法的扩展。通过功率和I类误差的计算和比较,探讨了这三种方法的优点。在两个块之间存在Ld的情况下,由于T-C忽略了两个块之间的相关性,因此T-C的I类误差高于T-H和T-G。对于三个统计量中的每一个,使用两个单倍型块的功率高于只使用一个单倍型块的功率。通过功率比较,我们注意到T-C的功率高于T-H的功率,T-H的功率高于T-G的功率。在两块之间没有LD的情况下,功率与T-H相似,但高于T-G。因此,我们主张在数据分析中使用T-H。在两个块之间存在LD的情况下,考虑了两个单倍型块之间的相关性,并且具有比T-G更低的I类错误和更高的功率。通过样本量计算,验证了该方法的可行性。
One way to perform linkage-disequilibrium (LD) mapping of genetic traits is to use single markers. Since dense marker maps-such as single- nucleotide polymorphism and high-resolution microsatellite maps-are available, it is natural and practical to generalize single- marker LD mapping to high-resolution haplotype or multiple-marker LD mapping. This article investigates high-resolution LD-mapping methods, for complex diseases, based on haplotype maps or microsatellite marker maps. The objective is to explore test statistics that combine information from haplotype blocks or multiple markers. Based on two coding methods, genotype coding and haplotype coding, Hotelling's T-2 statistics T-G and T-H are proposed to test the association between a disease locus and two haplotype blocks or two markers. The validity of the two T-2 statistics is proved by theoretical calculations. A statistic T-C, an extension of the traditional method of comparing haplotype frequencies, is introduced by simply adding the chi(2) test statistics of the two haplotype blocks together. The merit of the three methods is explored by calculation and comparison of power and of type I errors. In the presence of LD between the two blocks, the type I error of T-C is higher than that of T-H and T-G, since T-C ignores the correlation between the two blocks. For each of the three statistics, the power of using two haplotype blocks is higher than that of using only one haplotype block. By power comparison, we notice that T-C has higher power than that of T-H, and T-H has higher power than that of T-G. In the absence of LD between the two blocks, the power of is similar to that of T-H and higher than that of T-G. Hence, we advocate use of T-H in the data analysis. In the presence of LD between the two blocks, takes into account the correlation between the two haplotype blocks and has a lower type I error and higher power than T-G. Besides, the feasibility of the methods is shown by sample-size calculation.