Integrative analysis of multiple genomic variables using a hierarchical Bayesian model

Integrative analysis of multiple genomic variables using a hierarchical Bayesian model
复制标题

DOI:
10.1093/bioinformatics/btx356
复制
发表时间:
2017-10
期刊:
影响因子:
5.8
通讯作者:
Martin Schäfer;H. Klein;H. Schwender
Martin Schäfer;H. Klein;H. Schwender
中科院分区:
生物学3区
文献类型:
--
作者:
Martin Schäfer;H. Klein;H. Schwender

文献摘要

相似文献

在两种生物条件下,动机基因在几个基因组变量上显示出一致的差异,这对于揭开感兴趣表型背后的因果关系至关重要。检测这类基因在生物医学研究中很重要,例如在确定导致癌症发展的基因时。下一代测序研究中常见的小样本量是一个关键挑战,仍然只有很少的统计方法以综合的、基于模型的方式分析两个以上的基因组变量。在这里,我们提出了一种新的生物信息学方法来检测两种生物条件之间的一致性差异,在大量不同的测量中,例如不同的表观遗传标记或mRNA转录水平。结果我们提出了一个系数来量化基因在多个(两个以上)基因组变量中呈现一致变化的程度,当将呈现感兴趣的条件(例如癌症)的样本与参考组进行比较时。分层贝叶斯模型被用来评估基因水平上的不确定性,纳入了关于基因之间功能关系的信息。我们在包含RNA-seq基因转录子和多达四个ChIP-seq组蛋白修饰测量的不同数据集上演示了该方法。在分析多个基因组变量时,基于系数的排序和基于模型的推理都导致了候选基因合理的优先顺序。补充资料中的可用性和实现错误代码。联系m.schaefer@uni-duesseldorf.de补充信息补充数据可在BioInformation Online上获得。
Motivation Genes showing congruent differences in several genomic variables between two biological conditions are crucial to unravel causalities behind phenotypes of interest. Detecting such genes is important in biomedical research, e.g. when identifying genes responsible for cancer development. Small sample sizes common in next‐generation sequencing studies are a key challenge, and there are still only very few statistical methods to analyze more than two genomic variables in an integrative, model‐based way. Here, we present a novel bioinformatics approach to detect congruent differences between two biological conditions in a larger number of different measurements such as various epigenetic marks or mRNA transcript levels. Results We propose a coefficient quantifying the degree to which genes present consistent alterations in multiple (more than two) genomic variables when comparing samples presenting a condition of interest (e.g. cancer) to a reference group. A hierarchical Bayesian model is employed to assess uncertainty on a gene level, incorporating information on functional relationships between genes. We demonstrate the approach on different data sets containing RNA‐seq gene transcripton and up to four ChIP‐seq histone modification measurements. Both the coefficient‐based ranking and the inference based on the model lead to a plausible prioritizing of candidate genes when analyzing multiple genomic variables. Availability and implementation BUGS code in the Supplement. Contact m.schaefer@uni‐duesseldorf.de Supplementary information Supplementary data are available at Bioinformatics online.