Large-scale local surrogate modeling of stochastic simulation experiments

Large-scale local surrogate modeling of stochastic simulation experiments
复制标题

DOI:
10.1016/j.csda.2022.107537
复制
发表时间:
2021-09
期刊:
Comput. Stat. Data Anal.
影响因子:
--
通讯作者:
D. Cole;R. Gramacy;M. Ludkovski
D. Cole;R. Gramacy;M. Ludkovski
中科院分区:
其他
文献类型:
--
作者:
D. Cole;R. Gramacy;M. Ludkovski

文献摘要

被引文献

相似文献

大型计算机实验的高斯过程 (GP) 代理建模受到三次运行时间的限制,尤其是来自具有输入相关噪声的随机模拟的数据。降低计算复杂性的一种流行的解决方法涉及局部近似(例如 LAGP)。然而,LAGP 仅在确定性环境中经过审查。最近的一种变体利用诱导点 (LIGP) 来实现额外的稀疏性,在速度与准确性前沿上改进了 LAGP。作者表明,LIGP 相对于 LAGP 的另一个好处是,随机响应的(局部)块金估计更加自然,尤其是当设计包含大量复制时,这在尝试将信号与噪声分离时很常见。 Woodbury 恒等式在 LIGP 中从诱导点扩展到重复,仅在独特的设计位置方面提供有效的计算。这增加了无需额外触发器即可合并的本地数据量(即邻域大小),从而提高了统计效率。作者的 LIGP 升级的性能通过基准数据和现实世界的随机模拟实验(包括期权定价控制框架)进行了说明。结果表明,与现代替代方案相比,LIGP 对不同的数据维度和复制策略提供了更准确的预测和不确定性量化。
Gaussian process (GP) surrogate modeling for large computer experiments is limited by cubic runtimes, especially with data from stochastic simulations with input-dependent noise. A popular workaround to reduce computational complexity involves local approximation (e.g., LAGP). However, LAGP has only been vetted in deterministic settings. A recent variation utilizing inducing points (LIGP) for additional sparsity improves upon LAGP on the speed-vs-accuracy frontier. The authors show that another benefit of LIGP over LAGP is that (local) nugget estimation for stochastic responses is more natural, especially when designs contain substantial replication as is common when attempting to separate signal from noise. Woodbury identities, extended in LIGP from inducing points to replicates, afford efficient computation in terms of unique design locations only. This increases the amount of local data (i.e., the neighborhood size) that may be incorporated without additional flops, thereby enhancing statistical efficiency. Performance of the authors' LIGP upgrades is illustrated on benchmark data and real-world stochastic simulation experiments, including an options pricing control framework. Results indicate that LIGP provides more accurate prediction and uncertainty quantification for varying data dimension and replication strategies versus modern alternatives.