An efficient classification approach for large-scale mobile ubiquitous computing

An efficient classification approach for large-scale mobile ubiquitous computing
复制标题

大规模移动普适计算的高效分类方法

DOI:
10.1016/j.ins.2012.09.050
复制
发表时间:
2013-05-20
影响因子:
8.1
通讯作者:
Guo, Minyi
Guo, Minyi
中科院分区:
计算机科学1区
文献类型:
--
作者:
Tang, Feilong;You, Ilsun;Guo, Minyi

文献摘要

被引文献

相似文献

上下文分类是以用户为中心的普适计算的核心,其目标是基于表达的偏好和兴趣提供个性化服务。移动的普适计算(MUC)中存在大量的数据并且用户对上下文感知系统提出了很高的要求,因此上下文分类必须在计算方面是有效的和高效的。基于顺序最小优化(SMO)的SVMTorch被广泛应用于文本分类,但由于矩阵乘法效率低,分类速度慢,无法对多标签数据进行分类,因此在面向多用户上下文分析的SVMTorch中效率较低.首先,我们提出了一个半稀疏算法,以加快向量/矩阵乘法,这是基于SVMTorch的分类方法的核心。理论上,要将两个向量相乘(即,传统的SVMTorch算法需要O(m + n)时间,而我们的半稀疏算法只需要O(n)时间,其中n是训练向量中非零元素的个数。其次,我们扩展了传统的SVMTorch方法的功能,这是局限于单标签数据的分类,支持多类多标签分类。最后,将改进后的SVMTorch并行化,将半备用算法和功能扩展结合起来,以访问多核处理器和集群系统,进一步提高分类过程的有效性和效率。实验结果表明,我们提出的解决方案显着提高了传统的SVMTorch的性能和能力。实验结果表明,训练数据集和测试数据集越大,上下文分类的有效性和效率越高。在基于本文提出的解决方案开发的中文网页分类器中验证了这一结论。(C)2012 Elsevier Inc. All rights reserved.
Context classification is at the center of user-centric ubiquitous computing that targets the provision of personalized services based on expressed preferences and interests. Classification of context for Mobile Ubiquitous Computing (MUC), where there are high volumes of data and users place large demands on a context-aware system, must be effective and efficient in computational terms. The Sequential Minimal Optimization (SMO) based SVMTorch is widely used for text classification; it is however inefficient for MUC-oriented context analysis due to: (1) a low classification speed caused by inefficient matrix multiplication, and (2) the inability to classify multi-label data.In this paper, we propose an efficient classification approach to improve and extend the SVMTorch. Firstly, we propose a semi-sparse algorithm to speed up vector/matrix multiplication which lies at the core of the SVMTorch-based classification approaches. Theoretically, to multiply two vectors (i.e., a selected vector and a trained vector) with m and n non-zero elements respectively, the traditional SVMTorch needs O(m + n) time while our semi-sparse algorithm requires only O(n) time, where n is the number of non-zero elements in the trained vector. Secondly, we extend the functions of the traditional SVMTorch approach which is limited to the classification of single-label data, to support multi-class multi-label classification. Finally, we parallelize the improved SVMTorch which incorporates the semi-spares algorithm and function extensions to access multi-core processor and cluster systems to further improve the effectiveness and efficiency of the classification process. The experimental results demonstrate that our proposed solution significantly improves the performance and capability of the traditional SVMTorch. The results support the conclusion that the larger training and testing data sets are, the more improvement our solution brings to the effectiveness and efficiency of the context classification. This conclusion is verified in a Chinese web page classifier developed based on the solution presented in this paper. (C) 2012 Elsevier Inc. All rights reserved.