Predicting gene expression from genome wide protein binding profiles

Predicting gene expression from genome wide protein binding profiles
复制标题

DOI:
10.1016/j.neucom.2017.09.094
复制
发表时间:
2018-01-31
期刊:
影响因子:
6
通讯作者:
Wilson, Paul
Wilson, Paul
中科院分区:
计算机科学2区
文献类型:
--
作者:
Ferdous, Mohsina M.;Bao, Yanchun;Wilson, Paul

文献摘要

被引文献

相似文献

高通量技术,如染色质免疫沉淀(IP)和下一代测序(ChIP-seq),结合基因表达研究,使研究人员能够在全基因组范围内研究染色体相关蛋白的分布与基因转录调控之间的关系。一些综合分析的尝试已经确定了这两个过程之间的直接关系。然而,对监管事件的全面理解仍然难以捉摸。这部分是由于缺乏从ChIP-seq数据中检测结合区域的可靠分析方法。在本文中,我们应用了最近提出的马尔可夫随机场模型来检测不同生物条件和时间点下的富集结合区。该方法考虑了空间依赖性和IP效率,这在不同的实验之间可能会有很大的差异。我们进一步将富集的染色体结合区域定义为不同的基因组特征,如启动子、外显子、内含子和远端基因间,然后使用机器学习技术(包括神经网络、决策树和随机森林)研究了这些特征对基因表达活性的预测能力。从相同的生物样本中获得的ChIP-seq时间序列数据集包括六个蛋白质标记物和相关的微阵列数据,分析显示了有希望的结果,并确定了蛋白质谱和基因调控之间的生物学上合理的关系。(C) 2017年作者。Elsevier B.V.出版
High-throughput technologies such as chromatin immunoprecipitation (IP) followed by next generation sequencing (ChIP-seq) in combination with gene expression studies have enabled researchers to investigate relationships between the distribution of chromosome-associated proteins and the regulation of gene transcription on a genome-wide scale. Several attempts at integrative analyses have identified direct relationships between the two processes. However, a comprehensive understanding of the regulatory events remains elusive. This is in part due to the scarcity of robust analytical methods for the detection of binding regions from ChIP-seq data. In this paper, we have applied a recently proposed Markov random field model for the detection of enriched binding regions under different biological conditions and time points. The method accounts for spatial dependencies and IP efficiencies, which can vary significantly between different experiments. We further defined the enriched chromosomal binding regions as distinct genomic features, such as promoter, exon, intron, and distal intergenic, and then investigated how predictive each of these features are of gene expression activity using machine learning techniques, including neural networks, decision trees and random forest. The analysis of a ChIP-seq time-series dataset comprising six protein markers and associated microarray data, obtained from the same biological samples, shows promising results and identified biologically plausible relationships between the protein profiles and gene regulation. (C) 2017 The Authors. Published by Elsevier B.V.