UnTangle: Visual Mining for Data with Uncertain Multi-labels via Triangle Map

UnTangle: Visual Mining for Data with Uncertain Multi-labels via Triangle Map
复制标题

DOI:
10.1109/icdm.2014.24
复制
发表时间:
2014-12
期刊:
2014 IEEE International Conference on Data Mining
影响因子:
--
通讯作者:
Y. Lin;Nan Cao;D. Gotz;Lu Lu-Lu
Y. Lin;Nan Cao;D. Gotz;Lu Lu-Lu
中科院分区:
其他
文献类型:
--
作者:
Y. Lin;Nan Cao;D. Gotz;Lu Lu-Lu

文献摘要

被引文献

相似文献

具有多个不确定标签的数据在许多情况下是常见的。例如,一部电影可以与具有不同置信度的多个流派相关联,并且蛋白质序列可以被概率地分配给几个结构子类别。尽管它们无处不在,但将不确定的标签可视化的问题尚未得到充分解决。现有的方法通常要么丢弃不确定性信息,要么将数据映射到低维子空间,其中它们与多个标签的关联被模糊。本文提出了一种新的可视化挖掘技术Untangle,用于可视化不确定多标签。在我们提出的可视化中,数据项被放置在一个连接的三角形网络中,并为三角形顶点分配标签,以便附近的标签彼此更相关。基于项和标签之间的概率关联来确定数据项的位置。Untangle提供(A)自动标签放置算法和(B)允许用户控制针对不同视觉查询的标签定位的自适应交互机制。我们的工作提供了一种有效的方法来研究数据项与它们的不确定标签之间的关系,以及标签之间的关系,从而做出了独特的贡献。我们的用户研究表明,可视化有效地帮助用户发现紧急模式,并比较数据标签中不确定性信息的细微差别。
Data with multiple uncertain labels are common in many situations. For examples, a movie may be associated with multiple genres with different levels of confidence, and a protein sequence may be probabilistically assigned to several structural subcategories. Despite their ubiquity, the problem of visualizing uncertain labels has not been adequately addressed. Existing approaches often either discard the uncertainty information, or map the data to a low-dimensional subspace where their associations with multiple labels are obscured. In this paper, we propose a novel visual mining technique, UnTangle, for visualizing uncertain multi-labels. In our proposed visualization, data items are placed inside a web of connected triangles, with labels assigned to the triangle vertices such that nearby labels are more relevant to each other. The positions of the data items are determined based on the probabilistic associations between items and labels. UnTangle provides both (a) an automatic label placement algorithm, and (b) adaptive interaction mechanisms that allow users to control the label positioning for different visual queries. Our work makes a unique contribution by providing an effective way to investigate the relationship between data items and their uncertain labels, as well as the relationships among labels. Our user study suggests that the visualization effectively helps users discover emergent patterns and compare the nuances of uncertainty information in the data labels.