Forgetful Forests: Data Structures for Machine Learning on Streaming Data under Concept Drift

Forgetful Forests: Data Structures for Machine Learning on Streaming Data under Concept Drift
复制标题

DOI:
10.3390/a16060278
复制
发表时间:
2023-05
期刊:
影响因子:
2.3
通讯作者:
Zhehu Yuan;Yinqi Sun;D. Shasha
Zhehu Yuan;Yinqi Sun;D. Shasha
中科院分区:
--
文献类型:
--
作者:
Zhehu Yuan;Yinqi Sun;D. Shasha

文献摘要

相似文献

数据库和数据结构研究可以在许多方面提高机器学习性能。一种方法是设计更好的数据结构算法。本文结合使用增量计算以及顺序和概率过滤,使“健忘”基于树的学习算法,以科普流数据,遭受概念漂移。(当从输入到分类的功能映射随时间变化时,会发生概念漂移)。本文描述的遗忘算法实现了高性能,同时保持高质量的预测流数据。具体而言,该算法比最先进的增量算法快24倍,最多损失2%的准确度,或者至少快两倍而不损失任何准确度。这使得这样的结构适合于高容量流式传输应用。
Database and data structure research can improve machine learning performance in many ways. One way is to design better algorithms on data structures. This paper combines the use of incremental computation as well as sequential and probabilistic filtering to enable “forgetful” tree-based learning algorithms to cope with streaming data that suffers from concept drift. (Concept drift occurs when the functional mapping from input to classification changes over time). The forgetful algorithms described in this paper achieve high performance while maintaining high quality predictions on streaming data. Specifically, the algorithms are up to 24 times faster than state-of-the-art incremental algorithms with, at most, a 2% loss of accuracy, or are at least twice faster without any loss of accuracy. This makes such structures suitable for high volume streaming applications.