Knowledge graph embeddings for dealing with concept drift in machine learning

Knowledge graph embeddings for dealing with concept drift in machine learning
复制标题

用于处理机器学习中概念漂移的知识图谱嵌入

DOI:
10.1016/j.websem.2020.100625
复制
发表时间:
2021-01-22
影响因子:
2.5
通讯作者:
Chen, Huajun
Chen, Huajun
中科院分区:
计算机科学2区
文献类型:
--
作者:
Chen, Jiaoyan;Lecue, Freddy;Chen, Huajun

文献摘要

被引文献

相似文献

数据流学习主要用于从连续快速的数据记录中提取知识结构。由于数据在时间基础上不断发展,其基础知识受到许多挑战。概念漂移,作为流学习社区的核心挑战之一,被描述为数据的统计属性随时间的变化,导致大多数机器学习模型不太准确,因为随时间的变化是以不可预见的方式发生的。这尤其成问题,因为数据的演变可能导致知识的急剧变化。我们通过研究语义Web中数据流的语义表示(即本体流)来解决这个问题。这些流是用本体论词汇表注释的有序数据序列。特别是,我们利用本体流中编码的三个层次的知识来处理概念漂移:i)从流动力学中获得的新知识的存在,ii)知识变化和进化的意义,以及iii) (In)知识进化的一致性。这样的知识被编码为知识图嵌入,通过一种新颖表示的组合:蕴涵向量、蕴涵权重和一致性向量。我们在监督学习的分类任务上说明了我们的方法。本研究的主要贡献包括:(i)为流本体提供了一种有效的知识图嵌入方法;(ii)为处理概念漂移提供了一种集成知识图嵌入的通用一致预测框架。实验表明,我们的方法可以用现实世界的本体论流准确预测北京的空气质量和都柏林的公共汽车延误。(C) 2021 Elsevier B.V.版权所有
Data stream learning has been largely studied for extracting knowledge structures from continuous and rapid data records. As data is evolving on a temporal basis, its underlying knowledge is subject to many challenges. Concept drift,1 as one of core challenge from the stream learning community, is described as changes of statistical properties of the data over time, causing most of machine learning models to be less accurate as changes over time are in unforeseen ways. This is particularly problematic as the evolution of data could derive to dramatic change in knowledge. We address this problem by studying the semantic representation of data streams in the Semantic Web, i.e., ontology streams. Such streams are ordered sequences of data annotated with ontological vocabulary. In particular we exploit three levels of knowledge encoded in ontology streams to deal with concept drifts: i) existence of novel knowledge gained from stream dynamics, ii) significance of knowledge change and evolution, and iii) (in)consistency of knowledge evolution. Such knowledge is encoded as knowledge graph embeddings through a combination of novel representations: entailment vectors, entailment weights, and a consistency vector. We illustrate our approach on classification tasks of supervised learning. Key contributions of the study include: (i) an effective knowledge graph embedding approach for stream ontologies, and (ii) a generic consistent prediction framework with integrated knowledge graph embeddings for dealing with concept drifts. The experiments have shown that our approach provides accurate predictions towards air quality in Beijing and bus delay in Dublin with real world ontology streams. (C) 2021 Elsevier B.V. All rights reserved.