A generalized flow for multi-class and binary classification tasks: An Azure ML approach

A generalized flow for multi-class and binary classification tasks: An Azure ML approach
复制标题

多类和二元分类任务的通用流程:Azure ML 方法

DOI:
10.1109/bigdata.2015.7363944
复制
发表时间:
2015
期刊:
2015 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Sohini Roychowdhury
Sohini Roychowdhury
中科院分区:
--
文献类型:
--
作者:
Matthew Bihis;Sohini Roychowdhury

文献摘要

被引文献

相似文献

当今现实世界数据库的不断增长对单台计算机提出了计算挑战。另一方面,基于云的平台能够处理大量的信息操作任务,因此需要将其用于大型真实世界数据集计算。这项工作的重点是在基于云的计算平台中创建一个新的广义流:Microsoft Azure Machine Learning Studio(MAMLS),它接受多类和二进制分类数据集,并对其进行处理,以最大限度地提高整体分类准确性。首先,将每个数据集分别划分为训练数据集和测试数据集。然后,使用训练数据集估计线性和非线性分类模型参数。然后进行数据降维,以最大限度地提高分类精度。对于多类数据集,以数据为中心的信息被用来通过将多类分类减少到一系列分层二元分类任务来进一步提高整体分类精度。最后,在测试数据集上对优化后的分类模型的性能进行评估和评分。在3个公共数据集和一个本地数据集上,与现有的最先进的方法相比,对所提出的流的分类特性进行了比较评估。在3个公开数据集上,所提出的流实现了78-97.5%的分类准确率。此外,使用关于眼底图像中糖尿病视网膜病变病变的存在的信息创建的本地数据集导致85.3-95.7%的平均分类准确度,这高于现有方法。因此,所提出的广义流可以用于广泛的面向应用的“大数据集”。
The constant growth in the present day real-world databases pose computational challenges for a single computer. Cloud-based platforms, on the other hand, are capable of handling large volumes of information manipulation tasks, thereby necessitating their use for large real-world data set computations. This work focuses on creating a novel Generalized Flow within the cloud-based computing platform: Microsoft Azure Machine Learning Studio (MAMLS) that accepts multi-class and binary classification data sets alike and processes them to maximize the overall classification accuracy. First, each data set is split into training and testing data sets, respectively. Then, linear and nonlinear classification model parameters are estimated using the training data set. Data dimensionality reduction is then performed to maximize classification accuracy. For multi-class data sets, data-centric information is used to further improve overall classification accuracy by reducing the multi-class classification to a series of hierarchical binary classification tasks. Finally, the performance of optimized classification model thus achieved is evaluated and scored on the testing data set. The classification characteristics of the proposed flow are comparatively evaluated on 3 public data sets and a local data set with respect to existing state-of-the-art methods. On the 3 public data sets, the proposed flow achieves 78-97.5% classification accuracy. Also, the local data set, created using the information regarding presence of Diabetic Retinopathy lesions in fundus images, results in 85.3-95.7% average classification accuracy, which is higher than the existing methods. Thus, the proposed generalized flow can be useful for a wide range of application-oriented "big data sets".