Learning under Feature Drifts in Textual Streams

Learning under Feature Drifts in Textual Streams
复制标题

DOI:
10.1145/3269206.3271717
复制
发表时间:
2018-10
期刊:
Proceedings of the 27th ACM International Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Damianos P. Melidis;M. Spiliopoulou;Eirini Ntoutsi
Damianos P. Melidis;M. Spiliopoulou;Eirini Ntoutsi
中科院分区:
其他
文献类型:
--
作者:
Damianos P. Melidis;M. Spiliopoulou;Eirini Ntoutsi

文献摘要

被引文献

相似文献

如今会产生大量文本流,尤其是在 Twitter 和 Facebook 等社交网络中。由于讨论主题和用户对这些主题的看法随着时间的推移而发生巨大变化,这些流的数据分布发生变化,导致要学习的概念发生变化,这种现象称为概念漂移。一种尚未引起广泛关注的特殊类型的漂移是特征漂移,即与当前学习任务相关的特征的变化。在这项工作中,我们提出了一种处理文本流中特征漂移的方法。我们的方法集成了i)基于集成的机制,通过考虑不同的特征可能受到不同时间趋势的影响来准确预测下一个时间点的特征/单词值,以及ii)基于草图的特征空间维护机制,允许对流上的特征空间进行内存限制的维护。对来自情感分析、电子邮件偏好和垃圾邮件检测的文本流进行的实验表明,与基线相比,我们的方法取得了显着更好或具有竞争力的性能。
Huge amounts of textual streams are generated nowadays, especially in social networks like Twitter and Facebook. As the discussion topics and user opinions on those topics change drastically with time, those streams undergo changes in data distribution, leading to changes in the concept to be learned, a phenomenon called concept drift. One particular type of drift, that has not yet attracted a lot of attention is feature drift, i.e., changes in the features that are relevant for the learning task at hand. In this work, we propose an approach for handling feature drifts in textual streams. Our approach integrates i) an ensemble-based mechanism to accurately predict the feature/word values for the next time-point by taking into account the different features might be subject to different temporal trends and ii) a sketch-based feature space maintenance mechanism that allows for a memory-bounded maintenance of the feature space over the stream. Experiments with textual streams from the sentiment analysis, email preference and spam detection demonstrate that our approach achieves significantly better or competitive performance compared to baselines.