Augmenting Deep Learning with Relational Knowledge from Markov Logic Networks

Augmenting Deep Learning with Relational Knowledge from Markov Logic Networks
复制标题

DOI:
10.1109/bigdata50022.2020.9378055
复制
发表时间:
2020-12
期刊:
2020 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Mohammad Maminur Islam;Somdeb Sarkhel;D. Venugopal
Mohammad Maminur Islam;Somdeb Sarkhel;D. Venugopal
中科院分区:
其他
文献类型:
--
作者:
Mohammad Maminur Islam;Somdeb Sarkhel;D. Venugopal

文献摘要

被引文献

相似文献

神经符号学习将深度网络与符号知识相结合,可以帮助规范模型并控制过拟合。特别是,对于数据实例不独立的应用程序,领域知识可以用来指定关系依赖性,这可能很难纯粹从数据中推断出来。基于一阶逻辑的符号人工智能模型,如马尔可夫逻辑网络(MLN),被设计用于表示和推理不确定的背景知识。然而,这种模型中的学习和推理算法已知是缓慢和不准确的。在本文中,我们开发了一种新的模型,它结合了两个世界的最佳之处,即DNN的可扩展学习能力和MLN中指定的符号知识。为此,我们根据MLN知识库中编码的关系知识推断数据中的对称性,并训练卷积神经网络(CNN)来学习组合对称变量的内核。然而,通过这样做,我们被迫将关系数据拆分为独立的CNN训练实例,这可能会导致关系依赖性的丢失,从而为学习的模型增加噪声/不确定性。因此,我们不是学习单个模型,而是学习模型参数的分布。我们的实验表明,我们的模型在几个不同的问题领域优于纯MLN或纯DNN模型。
Neuro-symbolic learning, where deep networks are combined with symbolic knowledge can help regularize the model and control overfitting. In particular, for applications where data instances are not independent, domain knowledge can be used to specify relational dependencies which may be hard to infer purely from the data. Symbolic AI models such as Markov Logic networks (MLNs) which are based on first-order logic are designed to represent and reason with uncertain background knowledge. However learning and inference algorithms in such models is known to be slow and inaccurate. In this paper, we develop a novel model that combines the best of both worlds, namely, the scalable learning capabilities of DNNs and symbolic knowledge specified in MLNs. To do this, we infer symmetries in the data based on the relational knowledge encoded in an MLN knowledge base and train a Convolutional Neural Network (CNN) to learn kernels combining symmetrical variables. However, by doing this, we are forced to split the relational data into independent instances for CNN training which may result is a loss of relational dependencies adding noise/uncertainty to the learned model. Therefore, instead of a single model, we learn a distribution over the model parameters. Our experiments illustrate that our model outperforms purely-MLN or purely-DNN based models in several different problem domains.