Bayesian model selection for high-dimensional Ising models, with applications to educational data

Bayesian model selection for high-dimensional Ising models, with applications to educational data
复制标题

DOI:
10.1016/j.csda.2021.107325
复制
发表时间:
2019-11
期刊:
Comput. Stat. Data Anal.
影响因子:
--
通讯作者:
Jaewoo Park;Ick Hoon Jin-;M. Schweinberger
Jaewoo Park;Ick Hoon Jin-;M. Schweinberger
中科院分区:
其他
文献类型:
--
作者:
Jaewoo Park;Ick Hoon Jin-;M. Schweinberger

文献摘要

被引文献

相似文献

双难解后验分布出现在许多涉及离散和相关数据的统计学应用中,包括物理学、空间统计、机器学习、社会科学和其他领域。一个具体的例子是心理测量学,它采用了机器学习中的高维伊辛模型,以研究教育评估中二元项目反应之间的相互作用。为了从教育评估数据中估计高维的Ising模型,我们使用了1惩罚的节点逻辑回归。高维统计中的理论结果表明,只要满足一定的假设条件,1惩罚节点逻辑回归可以高概率地恢复真实的相互作用结构。这些假设在实践中很难验证并且可能被违背,并且量化估计的相互作用结构和参数估计的不确定性是具有挑战性的。我们提出了一种贝叶斯方法,该方法有助于量化相互作用结构和参数的不确定性,而不需要强假设,并且可以应用于具有数千个参数的Ising模型。我们通过模拟研究和应用于多达2,485个参数的小型和大型教育数据集,证明了所提出的贝叶斯方法与1惩罚节点逻辑回归相比的优势。除其他事项外,仿真研究表明,贝叶斯方法对由于省略协变量而导致的模型错误规范的鲁棒性比1惩罚的节点逻辑回归更强。
Doubly-intractable posterior distributions arise in many applications of statistics concerned with discrete and dependent data, including physics, spatial statistics, machine learning, the social sciences, and other fields. A specific example is psychometrics, which has adapted high-dimensional Ising models from machine learning, with a view to studying the interactions among binary item responses in educational assessments. To estimate high-dimensional Ising models from educational assessment data, ℓ 1-penalized nodewise logistic regressions have been used. Theoretical results in high-dimensional statistics show that ℓ 1-penalized nodewise logistic regressions can recover the true interaction structure with high probability, provided that certain assumptions are satisfied. Those assumptions are hard to verify in practice and may be violated, and quantifying the uncertainty about the estimated interaction structure and parameter estimators is challenging. We propose a Bayesian approach that helps quantify the uncertainty about the interaction structure and parameters without requiring strong assumptions, and can be applied to Ising models with thousands of parameters. We demonstrate the advantages of the proposed Bayesian approach compared with ℓ 1-penalized nodewise logistic regressions by simulation studies and applications to small and large educational data sets with up to 2,485 parameters. Among other things, the simulation studies suggest that the Bayesian approach is more robust against model misspecification due to omitted covariates than ℓ 1-penalized nodewise logistic regressions.