Coal-Miner: A Statistical Method for GWA Studies of Quantitative Traits with Complex Evolutionary Origins

Coal-Miner: A Statistical Method for GWA Studies of Quantitative Traits with Complex Evolutionary Origins
复制标题

DOI:
10.1145/3107411.3107490
复制
发表时间:
2017-08
期刊:
Proceedings of the 8th ACM International Conference on Bioinformatics, Computational Biology,and Health Informatics
影响因子:
--
通讯作者:
Hussein A. Hejase;N. V. Pol;G. Bonito;P. Edger;Kevin J. Liu
Hussein A. Hejase;N. V. Pol;G. Bonito;P. Edger;Kevin J. Liu
中科院分区:
其他
文献类型:
--
作者:
Hussein A. Hejase;N. V. Pol;G. Bonito;P. Edger;Kevin J. Liu

文献摘要

被引文献

相似文献

关联作图(AM)方法用于全基因组关联(GWA)研究,以检验基因型和表型数据之间的统计学显著关联。基因型和表型数据共享共同的进化起源-即,采样生物体的进化历史-引入协方差,其必须与GWA研究中主要感兴趣的生物功能的协方差区分开。已经引入了各种方法来执行AM,同时考虑样本相关性。然而,现有技术主要利用简化的假设,即样品相关性在整个基因组中是有效固定的。相反,群体遗传学理论和实证研究表明,样本相关性在基因组内的不同基因座之间可能差异很大。这种现象-称为局部系谱变异-在许多基因组数据集中经常遇到。需要新的AM方法来更好地解释基因组内样本相关性的局部变化。我们通过引入一种新的统计AM方法Coal-Miner来解决这一差距。Coal-Miner算法采用方法管道的形式。Coal-Miner的初始阶段寻求检测候选基因座,或含有pupillary相关标记的基因座。Coal-Miner的后续阶段使用具有多个效应的线性混合模型进行关联性检验,该模型解释了候选基因座内局部和整个基因组全局的样本相关性。使用合成和经验数据集,我们比较了统计功率和I型错误控制的煤矿工人对国家的最先进的AM方法。模拟条件反映了复杂性状的各种基因组结构,并结合了一系列的进化情景,每个情景都具有不同的进化过程,可以产生局部系谱变异。在我们研究的数据集中,我们发现与最先进的方法相比,Coal-Miner始终提供可比或通常更好的统计功效和I型错误控制。
Association mapping (AM) methods are used in genome-wide association (GWA) studies to test for statistically significant associations between genotypic and phenotypic data. The genotypic and phenotypic data share common evolutionary origins -- namely, the evolutionary history of sampled organisms -- introducing covariance which must be distinguished from the covariance due to biological function that is of primary interest in GWA studies. A variety of methods have been introduced to perform AM while accounting for sample relatedness. However, the state of the art predominantly utilizes the simplifying assumption that sample relatedness is effectively fixed across the genome. In contrast, population genetic theory and empirical studies have shown that sample relatedness can vary greatly across different loci within a genome. This phenomenon -- referred to as local genealogical variation -- is commonly encountered in many genomic datasets. New AM methods are needed to better account for local variation in sample relatedness within genomes. We address this gap by introducing Coal-Miner, a new statistical AM method. The Coal-Miner algorithm takes the form of a methodological pipeline. The initial stages of Coal-Miner seek to detect candidate loci, or loci which contain putatively associated markers. Subsequent stages of Coal-Miner perform test for association using a linear mixed model with multiple effects which account for sample relatedness locally within candidate loci and globally across the entire genome. Using synthetic and empirical datasets, we compare the statistical power and type I error control of Coal-Miner against state-of-the-art AM methods. The simulation conditions reflect a variety of genomic architectures for complex traits and incorporate a range of evolutionary scenarios, each with different evolutionary processes that can generate local genealogical variation. Across the datasets in our study, we find that Coal-Miner consistently offers comparable or typically better statistical power and type I error control compared to the state-of-the-art methods.