GPU-accelerated exhaustive search for third-order epistatic interactions in case-control studies

GPU-accelerated exhaustive search for third-order epistatic interactions in case-control studies
复制标题

DOI:
10.1016/j.jocs.2015.04.001
复制
发表时间:
2015-05-01
影响因子:
3.3
通讯作者:
Schmidt, Bertil
Schmidt, Bertil
中科院分区:
计算机科学3区
文献类型:
--
作者:
Gonzalez-Dominguez, Jorge;Schmidt, Bertil

文献摘要

被引文献

相似文献

近年来,从病例对照研究(如全基因组关联研究(GWAS))中发现与疾病密切相关的遗传标记组合的兴趣有所增加。检测上位性,即k个标记(k >= 2)之间的相互作用,是重要但耗时的操作,因为必须对测量的标记的每个k元组执行统计计算。对于k = 2,已经提出了有效的穷举方法,但是由于要计算的三元组的立方数,穷举三阶分析被认为是不切实际的。因此,大多数以前的方法应用解析法通过提前丢弃某些三元组来加速分析。不幸的是,这些工具可能无法检测到有趣的交互。我们提出了GPU3SNP,一个快速的GPU加速的工具,详尽地搜索一个给定的病例对照数据集的所有标记三元组之间的相互作用。我们的工具能够在合理的时间内分析具有数万个标记的输入数据集,这要归功于两个高效的CUDA内核和高效的工作负载分配技术。例如,一个由1000个个体测量的50,000个标记组成的数据集可以在不到22小时的时间内在一个具有4个NVIDIA GTX Titan板的计算节点上进行分析。(C)2015 Elsevier ay. All rights reserved.
Interest in discovering combinations of genetic markers from case-control studies, such as Genome Wide Association Studies (GWAS), that are strongly associated to diseases has increased in recent years. Detecting epistasis, i.e. interactions among k markers (k >= 2), is an important but time consuming operation since statistical computations have to be performed for each k-tuple of measured markers. Efficient exhaustive methods have been proposed for k = 2, but exhaustive third-order analyses are thought to be impractical due to the cubic number of triples to be computed. Thus, most previous approaches apply heuristics to accelerate the analysis by discarding certain triples in advance. Unfortunately, these tools can fail to detect interesting interactions. We present GPU3SNP, a fast GPU-accelerated tool to exhaustively search for interactions among all marker-triples of a given case-control dataset. Our tool is able to analyze an input dataset with tens of thousands of markers in reasonable time thanks to two efficient CUDA kernels and efficient workload distribution techniques. For instance, a dataset consisting of 50,000 markers measured from 1000 individuals can be analyzed in less than 22h on a single compute node with 4 NVIDIA GTX Titan boards. (C) 2015 Elsevier ay. All rights reserved.