Association mapping via regularized regression analysis of single-nucleotide-polymorphism haplotypes in variable-sized sliding windows

Association mapping via regularized regression analysis of single-nucleotide-polymorphism haplotypes in variable-sized sliding windows
复制标题

通过对大小可变的滑动窗口中的单核苷酸多态性单倍型进行正规化回归分析绘制关联图谱

DOI:
10.1086/513205
复制
发表时间:
2007-04-01
影响因子:
9.8
通讯作者:
Liu, Jian Jun
Liu, Jian Jun
中科院分区:
生物学1区
文献类型:
--
作者:
Li, Yi;Sung, Wing-Kin;Liu, Jian Jun

文献摘要

被引文献

相似文献

大规模的单倍型关联分析,特别是在全基因组水平上,仍然是一个非常具有挑战性的任务,没有最优的解决方案。在这项研究中,我们提出了一种新的单倍型关联分析方法,该方法基于可变大小的滑动窗口框架,并采用正则化回归分析来解决单倍型检验中的多自由度问题。我们的方法可以比现有的方法更有效地处理大量的单倍型关联分析。我们实施了一个程序,其中滑动窗口的最大大小由局部单倍型多样性和样本量决定,这是大规模单倍型分析的一个有吸引力的特征,如全基因组扫描,其中连锁不平衡模式预计会有很大的变化。我们使用模拟和实验数据,将我们的方法与其他三种方法的性能进行了比较-基于单核苷酸多态性的测试,单倍型的分支分析和变长马尔可夫链。通过分析在不同疾病模型下模拟的数据集,我们证明我们的方法始终优于其他三种方法,特别是当研究区域具有高单倍型多样性时。基于回归分析框架,我们的方法可以将其他风险因素信息整合到基于单倍型的关联分析中,这正在成为研究遗传和环境风险因素共同导致的常见疾病的越来越必要的步骤。
Large-scale haplotype association analysis, especially at the whole-genome level, is still a very challenging task without an optimal solution. In this study, we propose a new approach for haplotype association analysis that is based on a variable-sized sliding-window framework and employs regularized regression analysis to tackle the problem of multiple degrees of freedom in the haplotype test. Our method can handle a large number of haplotypes in association analyses more efficiently and effectively than do currently available approaches. We implement a procedure in which the maximum size of a sliding window is determined by local haplotype diversity and sample size, an attractive feature for large-scale haplotype analyses, such as a whole-genome scan, in which linkage disequilibrium patterns are expected to vary widely. We compare the performance of our method with that of three other methods - a test based on a single-nucleotide polymorphism, a cladistic analysis of haplotypes, and variable-length Markov chains - with use of both simulated and experimental data. By analyzing data sets simulated under different disease models, we demonstrate that our method consistently outperforms the other three methods, especially when the region under study has high haplotype diversity. Built on the regression analysis framework, our method can incorporate other risk-factor information into haplotype-based association analysis, which is becoming an increasingly necessary step for studying common disorders to which both genetic and environmental risk factors contribute.