Better together? Statistical learning in models made of modules

Better together? Statistical learning in models made of modules
复制标题

DOI:
--
复制
发表时间:
2017-08
期刊:
arXiv: Methodology
影响因子:
--
通讯作者:
P. Jacob;Lawrence M. Murray;C. Holmes;C. Robert
P. Jacob;Lawrence M. Murray;C. Holmes;C. Robert
中科院分区:
其他
文献类型:
--
作者:
P. Jacob;Lawrence M. Murray;C. Holmes;C. Robert

文献摘要

被引文献

相似文献

在现代应用中,统计学家面临着集成与推理、预测或决策问题相关的异构数据模式。在这种情况下,可以方便地使用图形模型通过一组连接的“模块”来表示统计依赖性,每个模块都与特定的数据模态相关,并在其开发中利用特定领域的专业知识。原则上,给定数据,传统的统计更新可以实现连贯的不确定性量化以及通过模块和跨模块的信息传播。然而,任何模块的错误指定都可能会以不可预测的方式影响其他模块的估计和更新。在各种设置中,特别是当某些模块比其他模块更受信任时,从业者倾向于避免使用完整模型进行学习,转而采用限制模块之间信息传播的方法,例如通过将传播限制为仅沿图边缘的特定方向。在本文中,我们研究了为什么在错误指定的设置中这些模块化方法可能比完整模型更好。我们提出了在模块化方法和全模型方法之间进行选择的原则标准。这个问题出现在许多应用环境中,包括大型随机动力系统、荟萃分析、流行病学模型、空气污染模型、药代动力学-药效学以及倾向评分的因果推断。
In modern applications, statisticians are faced with integrating heterogeneous data modalities relevant for an inference, prediction, or decision problem. In such circumstances, it is convenient to use a graphical model to represent the statistical dependencies, via a set of connected "modules", each relating to a specific data modality, and drawing on specific domain expertise in their development. In principle, given data, the conventional statistical update then allows for coherent uncertainty quantification and information propagation through and across the modules. However, misspecification of any module can contaminate the estimate and update of others, often in unpredictable ways. In various settings, particularly when certain modules are trusted more than others, practitioners have preferred to avoid learning with the full model in favor of approaches that restrict the information propagation between modules, for example by restricting propagation to only particular directions along the edges of the graph. In this article, we investigate why these modular approaches might be preferable to the full model in misspecified settings. We propose principled criteria to choose between modular and full-model approaches. The question arises in many applied settings, including large stochastic dynamical systems, meta-analysis, epidemiological models, air pollution models, pharmacokinetics-pharmacodynamics, and causal inference with propensity scores.