A benchmarking study of classification techniques for behavioral data

A benchmarking study of classification techniques for behavioral data
复制标题

DOI:
10.1007/s41060-019-00185-1
复制
发表时间:
2020-03-01
影响因子:
2.4
通讯作者:
Provost, Foster
Provost, Foster
中科院分区:
其他
文献类型:
--
作者:
De Cnudde, Sofie;Martens, David;Provost, Foster

文献摘要

被引文献

相似文献

越来越常见的大规模行为数据的预测能力已经被以前的研究所证明。这些数据通过人的动作和/或交互来捕获人的行为。它们的稀疏性和超高维度对最先进的分类技术提出了重大挑战。此外,没有以前的工作系统地探讨了选择的方法分类性能和计算费用之间的权衡。本文提供了一个贡献,在这个方向上通过基准研究。在41个细粒度行为数据集上比较了11个分类模型。统计性能比较丰富的学习曲线分析证明了两个重要的发现。首先,有一个固有的泛化性能与时间的权衡,使一个适当的分类器的选择依赖于计算约束和数据集的特点。良好正则化的逻辑回归实现了最好的AUC;然而,它需要最长的时间来训练。L2正则化的性能优于稀疏L1正则化。一个有吸引力的泛化/时间的权衡是通过一个基于相似性的技术。第二,虽然使用的数据集很大,但学习曲线结果表明,作为其高维性和稀疏性的直接结果,收集和分析更多数据具有重要价值。这一发现在实例和特征维度中都可以观察到,与传统数据的学习曲线研究形成对比。本研究的结果为研究人员和从业人员选择适当的分类技术,样本量和数据特征提供了指导,同时也为面对大型行为数据的可扩展算法设计提供了重点。
The predictive power of increasingly common large-scale, behavioral data has been demonstrated by previous research. Such data capture human behavior through the actions and/or interactions of people. Their sparsity and ultra-high dimensionality pose significant challenges to state-of-the-art classification techniques. Moreover, no prior work has systematically explored the choice of methods with respect to the trade-off between classification performance and computational expense. This paper provides a contribution in this direction through a benchmarking study. Eleven classification models are compared on forty-one fine-grained behavioral data sets. Statistical performance comparisons enriched with learning curve analyses demonstrate two important findings. First, there is an inherent generalization performance versus time trade-off, rendering the choice of an appropriate classifier dependent on computation constraints and data set characteristics. Well-regularized logistic regression achieves the best AUC; however, it takes the longest time to train. L2 regularization performs better than sparse L1 regularization. An attractive generalization/time trade-off is achieved by a similarity-based technique. Second, although the data sets used are large, the learning curve results illustrate that as a direct consequence of their high dimensionality and sparseness, significant value lies in collecting and analyzing even more data. This finding is observed both in the instance and in the feature dimensions, contrasting with learning curve studies on traditional data. The results of this study provide guidance for researchers and practitioners for the selection of appropriate classification techniques, sample sizes and data features, while also providing focus in scalable algorithm design in the face of large, behavioral data.