Interactive labelling of a multivariate dataset for supervised machine learning using linked visualisations, clustering, and active learning

Interactive labelling of a multivariate dataset for supervised machine learning using linked visualisations, clustering, and active learning
复制标题

DOI:
10.1016/j.visinf.2019.03.002
复制
发表时间:
2019-03-01
期刊:
影响因子:
3
通讯作者:
Schreck, Tobias
Schreck, Tobias
中科院分区:
计算机科学4区
文献类型:
--
作者:
Chegini, Mohammad;Bernard, Juergen;Schreck, Tobias

文献摘要

被引文献

相似文献

监督机器学习技术需要标记的多变量训练数据集。许多方法通过将机器学习算法与交互式可视化紧密结合来解决未标记数据集的问题。使用适当的技术,分析师可以在高度交互和迭代的机器学习过程中发挥积极作用,以标记数据集并创建有意义的分区。虽然这一原则已经实现了无监督,半监督,或监督机器学习任务,所有三种方法的组合仍然具有挑战性。在本文中,提出了一种可视化分析方法,结合了各种机器学习功能与四个链接的可视化视图,所有集成在mVis(多变量可视化)系统。可用的技术选项板允许分析师对多元数据集执行探索性数据分析,并将其划分为有意义的标记分区,从中可以构建分类器。在工作流程中,分析师可以在主动学习支持的半监督过程中标记有趣的模式或离群值。一旦数据集被交互式地标记,分析师就可以继续进行监督机器学习的工作流程,以评估后续分类器在多大程度上有效地学习了标记的训练数据集中表达的概念。使用一种称为自动维度选择的新技术,分析师与多变量数据集的维度的交互被用来引导机器学习算法。真实世界的足球数据集被用来展示mVis在一系列分析和标签任务中的效用,从初始标签到数据探索,聚类,分类和主动学习的迭代,以细化命名的分区,涉及最终产生适合于训练分类器的高质量标记的训练数据集。该工具使分析人员能够进行交互式可视化,包括散点图、平行坐标、记录的相似性图和分区的新相似性图。(C)2019浙江大学出版社.由爱思唯尔公司出版
Supervised machine learning techniques require labelled multivariate training datasets. Many approaches address the issue of unlabelled datasets by tightly coupling machine learning algorithms with interactive visualisations. Using appropriate techniques, analysts can play an active role in a highly interactive and iterative machine learning process to label the dataset and create meaningful partitions. While this principle has been implemented either for unsupervised, semi-supervised, or supervised machine learning tasks, the combination of all three methodologies remains challenging.In this paper, a visual analytics approach is presented, combining a variety of machine learning capabilities with four linked visualisation views, all integrated within the mVis (multivariate Visualiser) system. The available palette of techniques allows an analyst to perform exploratory data analysis on a multivariate dataset and divide it into meaningful labelled partitions, from which a classifier can be built. In the workflow, the analyst can label interesting patterns or outliers in a semi-supervised process supported by active learning. Once a dataset has been interactively labelled, the analyst can continue the workflow with supervised machine learning to assess to what degree the subsequent classifier has effectively learned the concepts expressed in the labelled training dataset. Using a novel technique called automatic dimension selection, interactions the analyst had with dimensions of the multivariate dataset are used to steer the machine learning algorithms.A real-world football dataset is used to show the utility of mVis for a series of analysis and labelling tasks, from initial labelling through iterations of data exploration, clustering, classification, and active learning to refine the named partitions, to finally producing a high-quality labelled training dataset suitable for training a classifier. The tool empowers the analyst with interactive visualisations including scatterplots, parallel coordinates, similarity maps for records, and a new similarity map for partitions. (C) 2019 Zhejiang University and Zhejiang University Press. Published by Elsevier B.V.