SparkINFERNO: a scalable high-throughput pipeline for inferring molecular mechanisms of non-coding genetic variants.

SparkINFERNO: a scalable high-throughput pipeline for inferring molecular mechanisms of non-coding genetic variants.
复制标题

SparkINFERNO:一个可扩展的高通量管道,用于推断非编码遗传变异的分子机制。

DOI:
10.1093/bioinformatics/btaa246
复制
发表时间:
2020
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Wang,Li-San
Wang,Li-San
中科院分区:
--
文献类型:
--
作者:
Kuksa,PavelP;Lee,Chien-Yueh;Amlie-Wolf,Alexandre;Gangadharan,Prabhakaran;Mlynarski,ElizabethE;Chou,Yi-Fan;Lin,Han-Jen;Issen,Heather;Greenfest-Allen,Emily;Valladares,Otto;Leung,YukYee;Wang,Li-San

文献摘要

相似文献

摘要我们报告了基于Spark的非编码遗传变异分子机制的推断(SparkINFERNO),这是一个可扩展的生物信息学管道,其特征是非编码全基因组关联研究(GWAS)的关联发现。SparkINFERNO优先考虑GWAs关联信号背后的因果变异,并报告相关的调节元件、组织背景和它们影响的可能的靶基因。为了实现这一目标,SparkINFERNO算法将Gwas汇总统计数据与涵盖400多个组织和细胞类型的增强子活性、转录因子结合、表达数量性状基因座和其他功能数据集的大规模功能基因组数据集整合在一起。可伸缩性是通过使用ApacheSpark和基于傻笑的基因组索引实现的底层API实现的。我们在大型GWAS上对SparkINFERNO进行了评估,结果表明SparkINFERNO的效率超过60倍,并且随着数据大小和计算资源量的增加而扩展。可用性和实施SparkINFERNO运行在具有APACHE Spark环境的群集或单台服务器上,并可在https://bitbucket.org/wanglab-upenn/SparkINFERNO或https://hub.docker.com/r/wanglab/spark-inferno.Contactlswang@pennmedicine.upenn.eduSupplementary信息网站上获得补充数据可从生物信息学在线获得
SummaryWe report Spark-based INFERence of the molecular mechanisms of NOn-coding genetic variants (SparkINFERNO), a scalable bioinformatics pipeline characterizing non-coding genome-wide association study (GWAS) association findings. SparkINFERNO prioritizes causal variants underlying GWAS association signals and reports relevant regulatory elements, tissue contexts and plausible target genes they affect. To achieve this, the SparkINFERNO algorithm integrates GWAS summary statistics with large-scale collection of functional genomics datasets spanning enhancer activity, transcription factor binding, expression quantitative trait loci and other functional datasets across more than 400 tissues and cell types. Scalability is achieved by an underlying API implemented using Apache Spark and Giggle-based genomic indexing. We evaluated SparkINFERNO on large GWASs and show that SparkINFERNO is more than 60 times efficient and scales with data size and amount of computational resources.Availability and implementationSparkINFERNO runs on clusters or a single server with Apache Spark environment, and is available at https://bitbucket.org/wanglab-upenn/SparkINFERNO or https://hub.docker.com/r/wanglab/spark-inferno.Contactlswang@pennmedicine.upenn.eduSupplementary informationSupplementary data are available atBioinformaticsonline