Learning a Tree-Structured Ising Model in Order to Make Predictions

Learning a Tree-Structured Ising Model in Order to Make Predictions
复制标题

DOI:
10.1214/19-aos1808
复制
发表时间:
2016-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Guy Bresler;Mina Karzand
Guy Bresler;Mina Karzand
中科院分区:
其他
文献类型:
--
作者:
Guy Bresler;Mina Karzand

文献摘要

相似文献

我们研究了从样本中学习树伊辛模型的问题,使得使用该模型进行的后续预测是准确的。在本文中考虑的预测任务是,预测的值的一个子集的变量给定值的一些其他子集的变量。几乎所有以前关于图形模型学习的工作都集中在恢复真正的底层图上。我们通过在给定大小的所有子集$\mathcal {S}$上取$P $和$Q$的边缘之间的总变化的最大值来定义分布$P $和$Q $之间的距离(“小集TV”或ssTV);这个距离捕获了感兴趣的预测任务的准确性。我们推导出非渐近界的样本数量需要得到一个分布(从同一类)与小ssTV相对于一个生成的样本。本文的主要信息之一是,需要的样本比恢复底层树少得多,这意味着使用错误的树可以进行准确的预测。
We study the problem of learning a tree Ising model from samples such that subsequent predictions made using the model are accurate. The prediction task considered in this paper is that of predicting the values of a subset of variables given values of some other subset of variables. Virtually all previous work on graphical model learning has focused on recovering the true underlying graph. We define a distance ("small set TV" or ssTV) between distributions $P$ and $Q$ by taking the maximum, over all subsets $\mathcal{S}$ of a given size, of the total variation between the marginals of $P$ and $Q$ on $\mathcal{S}$; this distance captures the accuracy of the prediction task of interest. We derive non-asymptotic bounds on the number of samples needed to get a distribution (from the same class) with small ssTV relative to the one generating the samples. One of the main messages of this paper is that far fewer samples are needed than for recovering the underlying tree, which means that accurate predictions are possible using the wrong tree.