课题基金 / 基金详情

项目摘要

项目成果

Vipul Periwal的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
With advances in single-cell techniques, collecting a large quantity of data has become more accessible and efficient. In contrast, the increased complexity of data has made it more challenging to unravel underlying biological mechanisms. Thus, it is critical to develop novel computational methods capable of dealing with such complexity and of providing some predictive deductions from the data. In this study, we present a novel method for the inference of a gene network using a new approach to least absolute deviation regression after using neural networks for imputing missing data. Fowlkes et al. (Cell, 2008) published a set of gene expressions measured from 6078 Drosophila blastoderm during six different time cohorts. Out of 95 genes and four proteins, only 27 of them had complete temporal information from all the cells, while the rest were measured only in a subset of cells. To impute the missing data, we trained and tested neural networks with one hidden layer on the complete 27 genes as predictors and the genes that were measured only in subsets of cells as targets. With the trained neural network, we imputed the missing gene expressions. To test the imputation methods performance, we arbitrarily selected three genes from the complete 27 genes and randomly removed time points from their gene profiles. Then, the missing values were imputed using the same method. The medians of the imputed values were compared to those of the observed values and showed negligible differences. We then developed a stable variation on least absolute deviation regression to infer a mechanistic model of the gene network that governs the discrete gene dynamics. The accuracy of this model was compared with a mechanistic model inferred with a standard least squared regression and with non-mechanistic neural networks with multiple hidden layers. All models were used to predict gene profiles given initial values. Since the regression methods have different cost functions, we compared the distributions of errors with two metrics, L1 norm and L2 norm. To understand training set size dependence of our results, we titrated the data set into smaller subsets to evaluate performance of all methods. In the limit of small sample size, the model inferred with our choice of regression performed better than the one inferred with least squared regression, but not as well as neural networks. However, in the limit of using the complete training set, the model inferred with the least absolute deviation regression showed higher predictive power than all other models. With the inferred mechanistic model, we made predictions on gene dynamics under various gene network perturbations, impossible in non- mechanistic models.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Adipocyte development and insulin resistance
Single Cell Data Analysis Algorithms
Liver regeneration after partial hepatectomy
Adipocyte development and insulin resistance
海外基金