SiGMoiD: A super-statistical generative model for binary data.

SiGMoiD: A super-statistical generative model for binary data.
复制标题

DOI:
10.1371/journal.pcbi.1009275
复制
发表时间:
2021-08
影响因子:
4.3
通讯作者:
Dixit PD
Dixit PD
中科院分区:
生物学2区
文献类型:
--
作者:
Zhao X;Plata G;Dixit PD

文献摘要

参考文献

被引文献

相似文献

在现代计算生物学中,人们对建立概率模型来描述大量共变二进制变量的集合非常感兴趣。然而,目前的方法来建立生成模型依赖于建模者的约束条件的识别和计算昂贵的推断时,变量的数量是大的(N~100)。在这里,我们用二进制数据的超统计生成模型(SiGMoiD)来解决这两个问题。SiGMoiD是一个基于最大熵的框架,我们将数据想象为来自超统计系统;给定样本中的单个二进制变量耦合到同一个“浴”,其强度变量因样本而异。重要的是,与modeler指定约束的标准最大熵方法不同,SiGMoiD算法直接从数据中推断出它们。由于这种约束的最佳选择,SiGMoiD允许我们对大量(N>1000)二进制变量的集合进行建模。最后,SiGMoiD提供了数据的降维描述,使我们能够识别相似数据点的聚类以及二进制变量。我们说明了SiGMoiD使用多个数据集跨越几个时间和长度尺度的多功能性。集体变化的二元变量在现代生物学中无处不在。鉴于这些系统的可能配置的数量通常远远超过可用样本的数量,生成模型已成为定量描述二进制数据的重要工具。构建生成模型的最新方法有几个概念上的限制。具体来说,它们依赖于建模者选择适合系统的约束,这在具有许多复杂交互的系统中可能具有挑战性。此外,当变量数量很大(N~100)时,推断它们的计算成本很高。为了解决这个问题,我们提出了最大熵方法的理论推广,使我们能够建模非常高维的数据,至少比目前可能的高一个数量级。这个框架将是一个显着的进步,在计算分析的协变二进制变量。
In modern computational biology, there is great interest in building probabilistic models to describe collections of a large number of co-varying binary variables. However, current approaches to build generative models rely on modelers’ identification of constraints and are computationally expensive to infer when the number of variables is large (N~100). Here, we address both these issues with Super-statistical Generative Model for binary Data (SiGMoiD). SiGMoiD is a maximum entropy-based framework where we imagine the data as arising from super-statistical system; individual binary variables in a given sample are coupled to the same ‘bath’ whose intensive variables vary from sample to sample. Importantly, unlike standard maximum entropy approaches where modeler specifies the constraints, the SiGMoiD algorithm infers them directly from the data. Due to this optimal choice of constraints, SiGMoiD allows us to model collections of a very large number (N>1000) of binary variables. Finally, SiGMoiD offers a reduced dimensional description of the data, allowing us to identify clusters of similar data points as well as binary variables. We illustrate the versatility of SiGMoiD using multiple datasets spanning several time- and length-scales. Collectively varying binary variables are ubiquitous in modern biology. Given that the number of possible configurations of these systems typically far exceeds the number of available samples, generative models have become an essential tool in quantitative descriptions of binary data. The state-of-the-art approaches to build generative models have several conceptual limitations. Specifically, they rely on the modeler choosing system-appropriate constraints, which can be challenging in systems with many complex interactions. Moreover, they are computationally expensive to infer when the number of variables is large (N~100). To address this issue, we propose a theoretical generalization of the maximum entropy approach that allows us to model very high dimensional data; at least an order of magnitude higher than what is currently possible. This framework will be a significant advancement in the computational analysis of covarying binary variables.
DOI: 10.1371/journal.pcbi.1000308
发表时间: 2009-03
影响因子: 4.3
作者:
Kumar VS;Maranas CD
通讯作者: Maranas CD
DOI: 10.1103/revmodphys.85.1115
发表时间: 2013-07-16
影响因子: 44.1
作者:
Presse, Steve;Ghosh, Kingshuk;Dill, Ken A.
通讯作者: Dill, Ken A.
DOI: 10.1038/s41467-020-18529-y
发表时间: 2020-09-21
影响因子: 16.6
作者:
Grilli J
通讯作者: Grilli J
DOI: 10.1103/revmodphys.88.035003
发表时间: 2016-07-26
影响因子: 44.1
作者:
Azaele, Sandro;Suweis, Samir;Maritan, Amos
通讯作者: Maritan, Amos
DOI: 10.1038/s41587-020-0660-7
发表时间: 2021-03
影响因子: 46.9
作者:
Martino C;Shenhav L;Marotz CA;Armstrong G;McDonald D;Vázquez-Baeza Y;Morton JT;Jiang L;Dominguez-Bello MG;Swafford AD;Halperin E;Knight R
通讯作者: Knight R