Entity-level stream classification: exploiting entity similarity to label the future observations referring to an entity

Entity-level stream classification: exploiting entity similarity to label the future observations referring to an entity
复制标题

DOI:
10.1007/s41060-019-00177-1
复制
发表时间:
2020-02-01
影响因子:
2.4
通讯作者:
Spiliopoulou, Myra
Spiliopoulou, Myra
中科院分区:
其他
文献类型:
--
作者:
Unnikrishnan, Vishnu;Beyer, Christian;Spiliopoulou, Myra

文献摘要

被引文献

相似文献

流分类算法传统上将到达的实例视为独立的。然而,在许多应用中,到达的示例可能取决于生成它们的“实体”,例如在产品评论中或在用户与应用服务器的交互中。在这项研究中,我们调查这种依赖性的潜力,通过分割的原始流的实例/“观察”到以实体为中心的子流,并将实体特定的信息到学习模型。我们提出了一个k-近邻启发流分类方法,在该方法中,到达观测的标签是通过利用属于这个实体的观测和类似于它的实体的知识来预测的。对于实体相似性的计算,我们考虑关于观测的知识和关于实体的知识,可能从一个域/特征空间不同的预测。为了区分这种知识转移有利于流分类的情况和实体上的知识无助于对观察结果进行分类的情况,我们还提出了一种基于使用k个随机实体(kRE)对子流进行随机采样的启发式方法。我们的学习场景不是完全监督的:在为每个实体的初始m个观测值获取标签后,我们假设没有额外的标签到达,并尝试从初始种子预测近期和远期观测值的标签。我们报告了我们从三个数据集的发现。
Stream classification algorithms traditionally treat arriving instances as independent. However, in many applications, the arriving examples may depend on the "entity" that generated them, e.g. in product reviews or in the interactions of users with an application server. In this study, we investigate the potential of this dependency by partitioning the original stream of instances/"observations" into entity-centric substreams and by incorporating entity-specific information into the learning model. We propose a k-nearest-neighbour-inspired stream classification approach, in which the label of an arriving observation is predicted by exploiting knowledge on the observations belonging to this entity and to entities similar to it. For the computation of entity similarity, we consider knowledge about the observations and knowledge about the entity, potentially from a domain/feature space different from that in which predictions are made. To distinguish between cases where this knowledge transfer is beneficial for stream classification and cases where the knowledge on the entities does not contribute to classifying the observations, we also propose a heuristic approach based on random sampling of substreams using k Random Entities (kRE). Our learning scenario is not fully supervised: after acquiring labels for the initial m observations of each entity, we assume that no additional labels arrive and attempt to predict the labels of near-future and far-future observations from that initial seed. We report on our findings from three datasets.