Bridging the gap between transcriptome and proteome measurements identifies post-translationally regulated genes

Bridging the gap between transcriptome and proteome measurements identifies post-translationally regulated genes
复制标题

DOI:
10.1093/bioinformatics/btt537
复制
发表时间:
2013-12-01
期刊:
影响因子:
5.8
通讯作者:
Niranjan, Mahesan
Niranjan, Mahesan
中科院分区:
生物学3区
文献类型:
--
作者:
Gunawardana, Yawwani;Niranjan, Mahesan

文献摘要

被引文献

相似文献

动机:尽管通过精确调节蛋白质浓度实现了许多动态细胞行为,但通过微阵列技术以及最近通过深度测序技术测量的信使RNA丰度被广泛用作蛋白质测量的替代物。尽管对于某些物种和在某些条件下,转录组和蛋白质组水平测量之间存在良好的相关性,但由于转录后和翻译后调节,这种相关性决不是普遍的,这两者在细胞中都是高度普遍的。在这里,我们寻求开发一种数据驱动的机器学习方法,以弥合这两个水平的高通量组学测量之间的差距酿酒酵母和部署模型以一种新的方式来揭示mRNA-蛋白质对的候选人翻译后regulation.Results:稀疏诱导回归(l(1)范数正则化)的特征选择的应用程序导致一组稳定的功能:即mRNA、核糖体占有率、核糖体密度、tRNA适应指数和密码子偏好,同时实现特征从37减少到5。与这些特征一起使用的线性预测器能够相当准确地预测蛋白质浓度(R-2 = 0:86)。蛋白质的浓度不能准确预测,作为离群值相对于预测,被证明有注释证据的翻译后修饰,显着超过随机子集的类似大小P
Motivation: Despite much dynamical cellular behaviour being achieved by accurate regulation of protein concentrations, messenger RNA abundances, measured by microarray technology, and more recently by deep sequencing techniques, are widely used as proxies for protein measurements. Although for some species and under some conditions, there is good correlation between transcriptome and proteome level measurements, such correlation is by no means universal due to post-transcriptional and post-translational regulation, both of which are highly prevalent in cells. Here, we seek to develop a data-driven machine learning approach to bridging the gap between these two levels of high-throughput omic measurements on Saccharomyces cerevisiae and deploy the model in a novel way to uncover mRNA-protein pairs that are candidates for post-translational regulation.Results: The application of feature selection by sparsity inducing regression (l(1) norm regularization) leads to a stable set of features: i.e. mRNA, ribosomal occupancy, ribosome density, tRNA adaptation index and codon bias while achieving a feature reduction from 37 to 5. A linear predictor used with these features is capable of predicting protein concentrations fairly accurately (R-2 = 0: 86). Proteins whose concentration cannot be predicted accurately, taken as outliers with respect to the predictor, are shown to have annotation evidence of post-translational modification, significantly more than random subsets of similar size P