Reverse‐Engineering Visualizations: Recovering Visual Encodings from Chart Images

Reverse‐Engineering Visualizations: Recovering Visual Encodings from Chart Images
复制标题

DOI:
10.1111/cgf.13193
复制
发表时间:
2017-06
影响因子:
2.5
通讯作者:
Jorge Poco;Jeffrey Heer
Jorge Poco;Jeffrey Heer
中科院分区:
计算机科学4区
文献类型:
--
作者:
Jorge Poco;Jeffrey Heer

文献摘要

被引文献

相似文献

我们研究了如何从图表图像中自动恢复视觉编码,主要使用推断的文本元素。我们提供了一个端到端管道,它以位图图像作为输入,并返回视觉编码规范作为输出。我们提出了一个文本分析管道,它检测图表中的文本元素,分类它们的角色(例如,图表标题,x轴标签,y轴标题等),并使用光学字符识别恢复文本内容。我们还训练了一个卷积神经网络用于标记类型分类。使用已识别的文本元素和图形标记类型,我们可以推断输入图表图像的编码规范。我们在三个图表语料库上评估了我们的技术:一组使用Vega生成的自动标记图表,来自Quartz新闻网站的图表,以及从学术论文中提取的图表。我们演示了跨各种输入图表类型的文本元素、标记类型和图表规范的准确自动推断。
We investigate how to automatically recover visual encodings from a chart image, primarily using inferred text elements. We contribute an end‐to‐end pipeline which takes a bitmap image as input and returns a visual encoding specification as output. We present a text analysis pipeline which detects text elements in a chart, classifies their role (e.g., chart title, x‐axis label, y‐axis title, etc.), and recovers the text content using optical character recognition. We also train a Convolutional Neural Network for mark type classification. Using the identified text elements and graphical mark type, we can then infer the encoding specification of an input chart image. We evaluate our techniques on three chart corpora: a set of automatically labeled charts generated using Vega, charts from the Quartz news website, and charts extracted from academic papers. We demonstrate accurate automatic inference of text elements, mark types, and chart specifications across a variety of input chart types.