Evaluating aggregate effects of rare and common variants in the 1000 Genomes Project exon sequencing data using latent variable structural equation modeling.

Evaluating aggregate effects of rare and common variants in the 1000 Genomes Project exon sequencing data using latent variable structural equation modeling.
复制标题

DOI:
10.1186/1753-6561-5-s9-s47
复制
发表时间:
2011-11-29
期刊:
影响因子:
--
通讯作者:
Zhang, Lx
Zhang, Lx
中科院分区:
其他
文献类型:
--
作者:
Nock, Nl;Zhang, Lx

文献摘要

被引文献

相似文献

评估稀有和常见变异的聚合效应的方法有限。因此,我们应用两阶段方法来评估1000个基因组计划数据中的聚合基因效应,这些数据包含来自7个群体的697个无关个体的24,487个单核苷酸多态(SNPs)。在第一阶段,我们使用单变量、多元回归模型确定了那些至少有一个SNP符合Bonferroni校正的潜在有趣基因(猪)。在第二阶段,我们使用结构方程模型的多元统计框架,通过将每个基因建模为一个由多个常见变量和罕见变量定义的潜在结构,来评估聚合猪对性状Q1的影响。在第一阶段,我们发现猪在随机选择的重复(137重复)和100个其他重复之间有明显的差异,但Flt1除外。在阶段1中,折叠罕见的变体减少了假阳性,但增加了假阴性。在第二阶段,我们建立了一个良好的模型,该模型包含了所有9个影响Q1的基因(Flt1、KDR、ArnT、ELAV4、Flt4、HIF1a、HIF3A、VEGFA、VEGFC),发现Flt1对Q1的影响最大(betastd=0.33±0.05)。使用Replate 137估计作为总体值,我们发现100个重复的参数(负载、路径、残差)及其标准误差的平均相对偏差不到5%。我们的潜变量扫描电子显微镜方法为模拟多个基因中罕见和常见变异的聚合效应提供了一个可行的框架,但在阶段1需要更好的方法来最小化类型I和类型II的错误。
Methods that can evaluate aggregate effects of rare and common variants are limited. Therefore, we applied a two-stage approach to evaluate aggregate gene effects in the 1000 Genomes Project data, which contain 24,487 single-nucleotide polymorphisms (SNPs) in 697 unrelated individuals from 7 populations. In stage 1, we identified potentially interesting genes (PIGs) as those having at least one SNP meeting Bonferroni correction using univariate, multiple regression models. In stage 2, we evaluate aggregate PIG effects on trait, Q1, by modeling each gene as a latent construct, which is defined by multiple common and rare variants, using the multivariate statistical framework of structural equation modeling (SEM). In stage 1, we found that PIGs varied markedly between a randomly selected replicate (replicate 137) and 100 other replicates, with the exception of FLT1. In stage 1, collapsing rare variants decreased false positives but increased false negatives. In stage 2, we developed a good-fitting SEM model that included all nine genes simulated to affect Q1 (FLT1, KDR, ARNT, ELAV4, FLT4, HIF1A, HIF3A, VEGFA, VEGFC) and found that FLT1 had the largest effect on Q1 (betastd = 0.33 ± 0.05). Using replicate 137 estimates as population values, we found that the mean relative bias in the parameters (loadings, paths, residuals) and their standard errors across 100 replicates was on average, less than 5%. Our latent variable SEM approach provides a viable framework for modeling aggregate effects of rare and common variants in multiple genes, but more elegant methods are needed in stage 1 to minimize type I and type II error.