Probabilistic Neural Data Fusion for Learning from an Arbitrary Number of Multi-fidelity Data Sets

Probabilistic Neural Data Fusion for Learning from an Arbitrary Number of Multi-fidelity Data Sets
复制标题

DOI:
10.1016/j.cma.2023.116207
复制
发表时间:
2023-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Carlos Mora;J. Eweis-Labolle;Tyler B. Johnson;Likith Gadde;R. Bostanabad
Carlos Mora;J. Eweis-Labolle;Tyler B. Johnson;Likith Gadde;R. Bostanabad
中科院分区:
其他
文献类型:
--
作者:
Carlos Mora;J. Eweis-Labolle;Tyler B. Johnson;Likith Gadde;R. Bostanabad

文献摘要

被引文献

相似文献

在工程和科学的许多应用中,分析人员可以同时访问多个数据源。在这种情况下,获取信息的总成本可以通过数据融合或多保真度(MF)建模来降低,其中利用廉价的低保真度(LF)源来减少对昂贵的高保真度(HF)数据的依赖。在本文中,我们采用神经网络(NN)的数据融合的情况下,数据是非常稀缺的,并从任意数量的来源与不同程度的保真度和成本。我们引入了一个独特的NN架构,MF建模转换成一个非线性流形学习问题。我们的NN架构反向学习非平凡的(例如,非加性和非分层)的LF源的偏差在一个可解释和可视化的流形,其中每个数据源是通过低维分布编码。该概率流形量化模型形式的不确定性,使得具有小偏差的LF源被编码为接近HF源。此外,我们赋予我们的NN的输出一个参数分布,不仅要量化任意的不确定性,而且要重新制定网络的损失函数的基础上严格适当的评分规则,提高鲁棒性和准确性看不见的HF数据。通过一组分析和工程实例,我们证明了我们的方法提供了很高的预测能力,同时量化了各种来源的不确定性。我们的代码和示例可以通过GitLab访问。
In many applications in engineering and sciences analysts have simultaneous access to multiple data sources. In such cases, the overall cost of acquiring information can be reduced via data fusion or multi-fidelity (MF) modeling where one leverages inexpensive low-fidelity (LF) sources to reduce the reliance on expensive high-fidelity (HF) data. In this paper, we employ neural networks (NNs) for data fusion in scenarios where data is very scarce and obtained from an arbitrary number of sources with varying levels of fidelity and cost. We introduce a unique NN architecture that converts MF modeling into a nonlinear manifold learning problem. Our NN architecture inversely learns non-trivial (e.g., non-additive and non-hierarchical) biases of the LF sources in an interpretable and visualizable manifold where each data source is encoded via a low-dimensional distribution. This probabilistic manifold quantifies model form uncertainties such that LF sources with small bias are encoded close to the HF source. Additionally, we endow the output of our NN with a parametric distribution not only to quantify aleatoric uncertainties, but also to reformulate the network’s loss function based on strictly proper scoring rules which improve robustness and accuracy on unseen HF data. Through a set of analytic and engineering examples, we demonstrate that our approach provides a high predictive power while quantifying various sources of uncertainty. Our codes and examples can be accessed via GitLab.