Implementing and evaluating a Gaussian mixture framework for identifying gene function from TnSeq data

Implementing and evaluating a Gaussian mixture framework for identifying gene function from TnSeq data
复制标题

实施和评估用于从 TnSeq 数据识别基因功能的高斯混合框架

DOI:
10.1142/9789813279827_0016
复制
发表时间:
2018
期刊:
Pacific Symposium on Biocomputing
影响因子:
--
通讯作者:
Tintle, Nathan
Tintle, Nathan
中科院分区:
--
文献类型:
--
作者:
Li, Kevin;Chen, Rachel;Lindsey, William;Best, Aaron;DeJongh, Matthew;Henry, Christopher;Tintle, Nathan

文献摘要

相似文献

微生物基因组测序的快速加速增加了了解细菌基因功能的机会。不幸的是,只有一小部分基因得到了研究。最近,TnSeq已被提出作为一种具有成本效益的,高度可靠的方法来预测基因功能,作为对基因组变化前后细胞适应性变化的响应。然而,主要的问题仍然是如何最好地确定是否观察到的数量变化的健身代表一个有意义的变化。为了解决这个问题,我们开发了一个高斯混合模型框架,用于对TnSeq实验中的基因功能进行分类。为了实现混合模型,我们提出了期望最大化算法和一个分层贝叶斯模型采样使用斯坦的汉密尔顿蒙特-卡罗采样器。我们将这些实现与当前TnSeq文献中使用的频率论方法进行比较。从模拟和大肠杆菌TnSeq实验产生的真实的数据,我们表明高斯混合框架的贝叶斯实现提供了最一致的分类结果。
The rapid acceleration of microbial genome sequencing increases opportunities to understand bacterial gene function. Unfortunately, only a small proportion of genes have been studied. Recently, TnSeq has been proposed as a cost-effective, highly reliable approach to predict gene functions as a response to changes in a cell’s fitness before-after genomic changes. However, major questions remain about how to best determine whether an observed quantitative change in fitness represents a meaningful change. To address the limitation, we develop a Gaussian mixture model framework for classifying gene function from TnSeq experiments. In order to implement the mixture model, we present the Expectation-Maximization algorithm and a hierarchical Bayesian model sampled using Stan’s Hamiltonian Monte-Carlo sampler. We compare these implementations against the frequentist method used in current TnSeq literature. From simulations and real data produced by E.coli TnSeq experiments, we show that the Bayesian implementation of the Gaussian mixture framework provides the most consistent classification results.