Collecting and Characterizing Natural Language Utterances for Specifying Data Visualizations

Collecting and Characterizing Natural Language Utterances for Specifying Data Visualizations
复制标题

DOI:
10.1145/3411764.3445400
复制
发表时间:
2021-05
期刊:
Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems
影响因子:
--
通讯作者:
Arjun Srinivasan;Nikhila Nyapathy;Bongshin Lee;S. Drucker;J. Stasko
Arjun Srinivasan;Nikhila Nyapathy;Bongshin Lee;S. Drucker;J. Stasko
中科院分区:
其他
文献类型:
--
作者:
Arjun Srinivasan;Nikhila Nyapathy;Bongshin Lee;S. Drucker;J. Stasko

文献摘要

被引文献

相似文献

用于数据可视化的自然语言界面(NLI)在学术研究和商业软件中越来越流行。然而,缺乏对人们如何通过自然语言指定可视化的经验理解。我们进行了一项在线研究(n = 102),向参与者展示了一系列可视化,并要求他们提供他们会提出的话语以生成显示的图表。从响应中,我们策划了一个893个话语的数据集,并根据(1)其措辞(例如命令,查询,问题)和(2)(2)它们所包含的信息(例如图表类型,数据聚合)来表征话语。为了帮助指导未来的研发,我们贡献了这个话语数据集,并将其应用于NLI的创建和基准进行可视化。
Natural language interfaces (NLIs) for data visualization are becoming increasingly popular both in academic research and in commercial software. Yet, there is a lack of empirical understanding of how people specify visualizations through natural language. We conducted an online study (N = 102), showing participants a series of visualizations and asking them to provide utterances they would pose to generate the displayed charts. From the responses, we curated a dataset of 893 utterances and characterized the utterances according to (1) their phrasing (e.g., commands, queries, questions) and (2) the information they contained (e.g., chart types, data aggregations). To help guide future research and development, we contribute this utterance dataset and discuss its applications toward the creation and benchmarking of NLIs for visualization.