Inferring couplings in networks across order-disorder phase transitions

Inferring couplings in networks across order-disorder phase transitions
复制标题

DOI:
10.1103/physrevresearch.4.023240
复制
发表时间:
2021-06
影响因子:
4.2
通讯作者:
Wave Ngampruetikorn;V. Sachdeva;J. Torrence;Jan Humplik;D. Schwab;S. Palmer
Wave Ngampruetikorn;V. Sachdeva;J. Torrence;Jan Humplik;D. Schwab;S. Palmer
中科院分区:
--
文献类型:
--
作者:
Wave Ngampruetikorn;V. Sachdeva;J. Torrence;Jan Humplik;D. Schwab;S. Palmer

文献摘要

相似文献

统计推断是许多科学工作的核心,但它如何工作仍然没有解决。要做到这一点,需要对统计模型、推理方法和数据结构之间的内在相互作用有定量的理解。为此,我们的直接耦合分析(DCA)-一个非常成功的方法,用于分析氨基酸序列数据推断成对的相互作用从随机图上的铁磁伊辛模型的样本的功效。我们的方法允许物理动机的探索定性不同的数据制度分离的相变。我们发现,推理质量强烈依赖于数据生成分布的性质:最佳精度发生在中间温度,从宏观秩序和热噪声的不利影响是最小的。重要的是,我们的研究结果表明,DCA并不总是优于其本地统计为基础的前辈,而DCA擅长在低温下,它变得不如简单的相关阈值在几乎所有的温度时,数据是有限的。我们的研究结果提供了深入了解DCA如此成功运作的机制,以及更广泛地说,推理如何与数据中的结构相互作用。
Statistical inference is central to many scientific endeavors, yet how it works remains unresolved. Answering this requires a quantitative understanding of the intrinsic interplay between statistical models, inference methods, and the structure in the data. To this end, we characterize the efficacy of direct coupling analysis (DCA) — a highly successful method for analyzing amino acid sequence data—in inferring pairwise interactions from samples of ferromagnetic Ising models on random graphs. Our approach allows for physically motivated exploration of qualitatively distinct data regimes separated by phase transitions. We show that inference quality depends strongly on the nature of data-generating distributions: optimal accuracy occurs at an intermediate temperature where the detrimental effects from macroscopic order and thermal noise are minimal. Importantly our results indicate that DCA does not always outperform its local-statistics-based predecessors; while DCA excels at low temperatures, it becomes inferior to simple correlation thresholding at virtually all temperatures when data are limited. Our findings offer insights into the regime in which DCA operates so successfully, and more broadly, how inference interacts with the structure in the data.