SiGMoiD: A super-statistical generative model for binary data.
SiGMoiD: A super-statistical generative model for binary data.
复制标题
DOI:
10.1371/journal.pcbi.1009275
复制
发表时间:
2021-08
影响因子:
4.3
通讯作者:
Dixit PD
中科院分区:
文献类型:
--
作者:
Zhao X;Plata G;Dixit PD
In modern computational biology, there is great interest in building probabilistic models to describe collections of a large number of co-varying binary variables. However, current approaches to build generative models rely on modelers’ identification of constraints and are computationally expensive to infer when the number of variables is large (N~100). Here, we address both these issues with Super-statistical Generative Model for binary Data (SiGMoiD). SiGMoiD is a maximum entropy-based framework where we imagine the data as arising from super-statistical system; individual binary variables in a given sample are coupled to the same ‘bath’ whose intensive variables vary from sample to sample. Importantly, unlike standard maximum entropy approaches where modeler specifies the constraints, the SiGMoiD algorithm infers them directly from the data. Due to this optimal choice of constraints, SiGMoiD allows us to model collections of a very large number (N>1000) of binary variables. Finally, SiGMoiD offers a reduced dimensional description of the data, allowing us to identify clusters of similar data points as well as binary variables. We illustrate the versatility of SiGMoiD using multiple datasets spanning several time- and length-scales. Collectively varying binary variables are ubiquitous in modern biology. Given that the number of possible configurations of these systems typically far exceeds the number of available samples, generative models have become an essential tool in quantitative descriptions of binary data. The state-of-the-art approaches to build generative models have several conceptual limitations. Specifically, they rely on the modeler choosing system-appropriate constraints, which can be challenging in systems with many complex interactions. Moreover, they are computationally expensive to infer when the number of variables is large (N~100). To address this issue, we propose a theoretical generalization of the maximum entropy approach that allows us to model very high dimensional data; at least an order of magnitude higher than what is currently possible. This framework will be a significant advancement in the computational analysis of covarying binary variables.
登录
查看更多内容
影响因子:
4.3
作者:
Kumar VS;Maranas CD
通讯作者:
Maranas CD
影响因子:
44.1
作者:
Presse, Steve;Ghosh, Kingshuk;Dill, Ken A.
通讯作者:
Dill, Ken A.
影响因子:
16.6
作者:
Grilli J
通讯作者:
Grilli J
影响因子:
44.1
作者:
Azaele, Sandro;Suweis, Samir;Maritan, Amos
通讯作者:
Maritan, Amos
影响因子:
46.9
作者:
Martino C;Shenhav L;Marotz CA;Armstrong G;McDonald D;Vázquez-Baeza Y;Morton JT;Jiang L;Dominguez-Bello MG;Swafford AD;Halperin E;Knight R
通讯作者:
Knight R