Powerful differential expression analysis incorporating network topology for next-generation sequencing data

Powerful differential expression analysis incorporating network topology for next-generation sequencing data
复制标题

DOI:
10.1093/bioinformatics/btw833
复制
发表时间:
2017-05-15
期刊:
影响因子:
5.8
通讯作者:
Salim, Agus
Salim, Agus
中科院分区:
生物学3区
文献类型:
--
作者:
Dona, Malathi S. I.;Prendergast, Luke A.;Salim, Agus

文献摘要

被引文献

相似文献

动机:RNA-seq已成为查询转录组的首选技术。然而,大多数RNA-seq差异表达(DE)分析方法没有利用生物网络的先验知识来检测DE基因。随着生物网络数据库的可用性和质量的提高,可以利用这种先验知识的方法是必要的,这将为生物学家在分析RNA-seq数据时提供一个可行的、更强大的选择。结果:我们提出了一种三态马尔可夫随机场(MRF)方法,该方法利用已知的生物学途径和相互作用来提高灵敏度和特异性,从而降低从RNA-seq数据中检测差异表达基因的错误发现率(FDRs)。该方法需要规范化计数数据(例如Fragments或Reads Per Kilobase of transcript Per Million mapped Reads (FPKM/RPKM)格式)作为输入,它是在一个R包pathDESeq中实现的,可以从Github获得。仿真研究表明,该方法在不同样本量下都优于双态MRF模型。此外,对于类似的FDR,它具有比DESeq, EBSeq, edgeR和NOISeq更好的灵敏度。当应用于结直肠癌和肝细胞癌研究的真实数据集时,该方法还分别选择了更多的顶级基因本体术语和KEGG通路术语。总的来说,这些发现清楚地强调了我们的方法相对于不利用生物网络先验知识的现有方法的力量。补充信息:补充数据可在生物信息学在线获取。
Motivation: RNA-seq has become the technology of choice for interrogating the transcriptome. However, most methods for RNA-seq differential expression (DE) analysis do not utilize prior knowledge of biological networks to detect DE genes. With the increased availability and quality of biological network databases, methods that can utilize this prior knowledge are needed and will offer biologists with a viable, more powerful alternative when analyzing RNA-seq data.Results: We propose a three-state Markov Random Field (MRF) method that utilizes known biological pathways and interaction to improve sensitivity and specificity and therefore reducing false discovery rates (FDRs) when detecting differentially expressed genes from RNA-seq data. The method requires normalized count data (e.g. in Fragments or Reads Per Kilobase of transcript per Million mapped reads (FPKM/RPKM) format) as its input and it is implemented in an R package pathDESeq available from Github. Simulation studies demonstrate that our method outperforms the two-state MRF model for various sample sizes. Furthermore, for a comparable FDR, it has better sensitivity than DESeq, EBSeq, edgeR and NOISeq. The proposed method also picks more top Gene Ontology terms and KEGG pathways terms when applied to real dataset from colorectal cancer and hepatocellular carcinoma studies, respectively. Overall, these findings clearly highlight the power of our method relative to the existing methods that do not utilize prior knowledge of biological network.Supplementary information: Supplementary data are available at Bioinformatics online.