课题基金 / 基金详情

项目摘要

项目成果

Vipul Periwal的其他基金

相似基金

相关文献

中文摘要
翻译
随着单细胞技术的进步,收集大量数据变得更加容易和有效。相比之下,数据的复杂性增加了揭示潜在生物机制的挑战性。因此,它是至关重要的,以开发新的计算方法,能够处理这种复杂性,并提供一些预测的数据推断。在这项研究中,我们提出了一种新的方法,使用一种新的方法,最小绝对偏差回归后,使用神经网络插补缺失数据的基因网络的推理。 Fowlkes等人(Cell,2008)发表了一组在6个不同时间组期间从6078个果蝇胚盘测量的基因表达。在95个基因和4个蛋白质中,只有27个具有来自所有细胞的完整时间信息,而其余的仅在一个细胞子集中测量。为了填补缺失的数据,我们训练和测试了神经网络,其中一个隐藏层是完整的27个基因作为预测因子,而仅在细胞子集中测量的基因作为目标。通过训练好的神经网络,我们估算了缺失的基因表达。为了测试插补方法的性能,我们从完整的27个基因中任意选择了3个基因,并从其基因谱中随机删除了时间点。然后,使用相同的方法插补缺失值。将插补值的中位数与观察值的中位数进行比较,结果显示差异可忽略不计。 然后,我们开发了一个稳定的最小绝对偏差回归的变化来推断一个机制模型的基因网络,管理离散的基因动态。该模型的准确性进行了比较与一个机械模型推断与标准的最小二乘回归和非机械神经网络与多个隐藏层。所有模型都用于预测给定初始值的基因谱。由于回归方法具有不同的成本函数,因此我们用两个度量(L1范数和L2范数)比较了误差的分布。为了理解我们的结果对训练集大小的依赖性,我们将数据集滴定为较小的子集,以评估所有方法的性能。在小样本量的限制下,我们选择的回归推断的模型比最小二乘回归推断的模型表现得更好,但不如神经网络。然而,在使用完整训练集的限制下,用最小绝对偏差回归推断的模型显示出比所有其他模型更高的预测能力。利用所推导的机制模型,我们对各种基因网络扰动下的基因动态进行了预测,这在非机制模型中是不可能的。
英文摘要
With advances in single-cell techniques, collecting a large quantity of data has become more accessible and efficient. In contrast, the increased complexity of data has made it more challenging to unravel underlying biological mechanisms. Thus, it is critical to develop novel computational methods capable of dealing with such complexity and of providing some predictive deductions from the data. In this study, we present a novel method for the inference of a gene network using a new approach to least absolute deviation regression after using neural networks for imputing missing data. Fowlkes et al. (Cell, 2008) published a set of gene expressions measured from 6078 Drosophila blastoderm during six different time cohorts. Out of 95 genes and four proteins, only 27 of them had complete temporal information from all the cells, while the rest were measured only in a subset of cells. To impute the missing data, we trained and tested neural networks with one hidden layer on the complete 27 genes as predictors and the genes that were measured only in subsets of cells as targets. With the trained neural network, we imputed the missing gene expressions. To test the imputation methods performance, we arbitrarily selected three genes from the complete 27 genes and randomly removed time points from their gene profiles. Then, the missing values were imputed using the same method. The medians of the imputed values were compared to those of the observed values and showed negligible differences. We then developed a stable variation on least absolute deviation regression to infer a mechanistic model of the gene network that governs the discrete gene dynamics. The accuracy of this model was compared with a mechanistic model inferred with a standard least squared regression and with non-mechanistic neural networks with multiple hidden layers. All models were used to predict gene profiles given initial values. Since the regression methods have different cost functions, we compared the distributions of errors with two metrics, L1 norm and L2 norm. To understand training set size dependence of our results, we titrated the data set into smaller subsets to evaluate performance of all methods. In the limit of small sample size, the model inferred with our choice of regression performed better than the one inferred with least squared regression, but not as well as neural networks. However, in the limit of using the complete training set, the model inferred with the least absolute deviation regression showed higher predictive power than all other models. With the inferred mechanistic model, we made predictions on gene dynamics under various gene network perturbations, impossible in non- mechanistic models.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Adipocyte development and insulin resistance
Single Cell Data Analysis Algorithms
Liver regeneration after partial hepatectomy
Adipocyte development and insulin resistance
海外基金