Estimating replicate time shifts using Gaussian process regression

Estimating replicate time shifts using Gaussian process regression
复制标题

DOI:
10.1093/bioinformatics/btq022
复制
发表时间:
2010-03-15
期刊:
影响因子:
5.8
通讯作者:
Ihler, Alexander
Ihler, Alexander
中科院分区:
生物学3区
文献类型:
--
作者:
Liu, Qiang;Lin, Kevin K.;Ihler, Alexander

文献摘要

被引文献

相似文献

动机:时程基因表达数据集为生物过程的动态方面提供了重要的见解,如昼夜节律,细胞周期和器官发育。在典型的微阵列时程实验中,在每个时间点从多个重复样品获得测量值。从实验观察中准确地恢复基因表达模式是具有挑战性的测量噪声和重复的发展速度之间的变化。关于这个主题的先前工作集中在假设复制时间是同步的表达模式的推断上。我们开发了一种统计方法,同时推断(i)每个基因的潜在(隐藏)表达谱,以及(ii)每个单独重复的生物时间。我们的方法是基于高斯过程回归(GPR)结合一个概率模型,占每个replication.Results的生物发育时间的不确定性:我们应用GPR与不确定的测量时间的mRNA表达的微阵列数据集的毛发生长周期在小鼠背部皮肤,预测每个重复的轮廓形状和生物时间。预测的时移与独立获得的相对发育的形态估计显示出高度一致性。我们还表明,该方法系统地减少了预测误差的样本外的数据,显着降低了交叉验证研究中的均方误差。
Motivation: Time-course gene expression datasets provide important insights into dynamic aspects of biological processes, such as circadian rhythms, cell cycle and organ development. In a typical microarray time-course experiment, measurements are obtained at each time point from multiple replicate samples. Accurately recovering the gene expression patterns from experimental observations is made challenging by both measurement noise and variation among replicates' rates of development. Prior work on this topic has focused on inference of expression patterns assuming that the replicate times are synchronized. We develop a statistical approach that simultaneously infers both (i) the underlying (hidden) expression profile for each gene, as well as (ii) the biological time for each individual replicate. Our approach is based on Gaussian process regression (GPR) combined with a probabilistic model that accounts for uncertainty about the biological development time of each replicate.Results: We apply GPR with uncertain measurement times to a microarray dataset of mRNA expression for the hair-growth cycle in mouse back skin, predicting both profile shapes and biological times for each replicate. The predicted time shifts show high consistency with independently obtained morphological estimates of relative development. We also show that the method systematically reduces prediction error on out-of-sample data, significantly reducing the mean squared error in a cross-validation study.