Family-based association studies for next-generation sequencing.

Family-based association studies for next-generation sequencing.
复制标题

DOI:
10.1016/j.ajhg.2012.04.022
复制
发表时间:
2012-06
影响因子:
9.8
通讯作者:
Yun Zhu;M. Xiong
Yun Zhu;M. Xiong
中科院分区:
生物学1区
文献类型:
--
作者:
Yun Zhu;M. Xiong

文献摘要

被引文献

相似文献

一个人的疾病风险是由两种常见变异的复合作用决定的,一种是从遥远的祖先遗传而来的,在人群中分离,另一种是从最近的祖先遗传而来的罕见变异,主要在谱系中分离。下一代测序(NGS)技术生成的高维数据允许对遗传变异进行近乎完整的评估。尽管NGS技术前景光明,但也存在显著的局限性:错误率高,稀有变异富集,大部分缺失值,以及大多数当前分析方法是为基于人群的关联研究而设计的。为了应对NGS提出的分析挑战,我们提出了一个基于序列的关联研究的一般框架,该框架可以使用从任何人口结构中采样的各种类型的家庭和无关个体数据,以及一个通用的程序,可以转换任何基于人口的关联检验统计量,用于基于家庭的关联检验。我们开发了基于家族的功能性主成分分析(FPCA),有或没有平滑,广义T2,结合多元和崩溃(CMC)的方法,和单标记关联检验统计。通过密集的模拟,我们证明了基于家族的平滑FPCA(SFPCA)具有正确的I型错误率,并且能够更有效地检测(1)常见变异,(2)罕见变异,(3)常见和罕见变异,以及(4)与其他基于群体或基于家族的关联分析方法具有相反方向的影响的变异。建议的统计量被应用到两个数据集的系谱结构。结果表明,平滑后的FPCA具有比其他统计量小得多的p值。
An individual's disease risk is determined by the compounded action of both common variants, inherited from remote ancestors, that segregated within the population and rare variants, inherited from recent ancestors, that segregated mainly within pedigrees. Next-generation sequencing (NGS) technologies generate high-dimensional data that allow a nearly complete evaluation of genetic variation. Despite their promise, NGS technologies also suffer from remarkable limitations: high error rates, enrichment of rare variants, and a large proportion of missing values, as well as the fact that most current analytical methods are designed for population-based association studies. To meet the analytical challenges raised by NGS, we propose a general framework for sequence-based association studies that can use various types of family and unrelated-individual data sampled from any population structure and a universal procedure that can transform any population-based association test statistic for use in family-based association tests. We develop family-based functional principal-component analysis (FPCA) with or without smoothing, a generalizedT2, combined multivariate and collapsing (CMC) method, and single-marker association test statistics. Through intensive simulations, we demonstrate that the family-based smoothed FPCA (SFPCA) has the correct type I error rates and much more power to detect association of (1) common variants, (2) rare variants, (3) both common and rare variants, and (4) variants with opposite directions of effect from other population-based or family-based association analysis methods. The proposed statistics are applied to two data sets with pedigree structures. The results show that the smoothed FPCA has a much smaller p value than other statistics.