VASTA: a vision and language-assisted smartphone task automation system

VASTA: a vision and language-assisted smartphone task automation system
复制标题

VASTA:视觉和语言辅助的智能手机任务自动化系统

DOI:
--
复制
发表时间:
2019
期刊:
International Conference on Intelligent User Interfaces
影响因子:
--
通讯作者:
Iqbal Mohomed
Iqbal Mohomed
中科院分区:
--
文献类型:
--
作者:
A. R. Sereshkeh;Gary Leung;K. Perumal;Caleb Phillips;Minfan Zhang;A. Fazly;Iqbal Mohomed

文献摘要

参考文献

被引文献

相似文献

我们推出了 VASTA,这是一种用于智能手机任务自动化的新型视觉和语言辅助演示编程 (PBD) 系统。开发强大的 PBD 自动化系统需要克服三个关键挑战:首先,如何使特定演示对用户界面 (UI) 元素中的位置和视觉变化具有鲁棒性;其次,如何识别自动化参数的变化,使演示尽可能具有普适性;第三,如何从用户话语中识别用户希望执行什么自动化。为了解决第一个挑战,VASTA 利用最先进的计算机视觉技术(包括对象检测和光学字符识别)来准确标记用户演示的交互,而不依赖于底层 UI 结构。为了解决第二个和第三个挑战,VASTA 利用先进的自然语言理解算法来分析用户话语,以触发 VASTA 自动化脚本,并确定泛化的自动化参数。我们进行了一项初始用户研究,展示了 VASTA 在对用户话语进行聚类、了解自动化参数的变化、检测所需的 UI 元素以及最重要的是自动化各种任务方面的有效性。该系统的演示视频可在此处获取:http://y2u.be/kr2xE-FixjI。
We present VASTA, a novel vision and language-assisted Programming By Demonstration (PBD) system for smartphone task automation. Development of a robust PBD automation system requires overcoming three key challenges: first, how to make a particular demonstration robust to positional and visual changes in the user interface (UI) elements; secondly, how to recognize changes in the automation parameters to make the demonstration as generalizable as possible; and thirdly, how to recognize from the user utterance what automation the user wishes to carry out. To address the first challenge, VASTA leverages state-of-the-art computer vision techniques, including object detection and optical character recognition, to accurately label interactions demonstrated by a user, without relying on the underlying UI structures. To address the second and third challenges, VASTA takes advantage of advanced natural language understanding algorithms for analyzing the user utterance to trigger the VASTA automation scripts, and to determine the automation parameters for generalization. We run an initial user study that demonstrates the effectiveness of VASTA at clustering user utterances, understanding changes in the automation parameters, detecting desired UI elements, and, most importantly, automating various tasks. A demo video of the system is available here: http://y2u.be/kr2xE-FixjI.
通过自然语言指令和 GUI 演示进行交互式任务和概念学习
DOI: --
发表时间: 2020
期刊: The AAAI-20 Workshop on Intelligent Process Automation (IPA-20
影响因子: --
作者:
Li, Toby Jia-Jun;Radensky, Marissa;Jia, Justin;Singarajah, Kirielle;Mitchell, Tom M.;Myers, Brad A.
通讯作者: Myers, Brad A.
DOI: 10.1145/3242587.3242650
发表时间: 2018-10
期刊: Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology
影响因子: --
作者:
Thomas F. Liu;Mark Craft;Jason Situ;Ersin Yumer;R. Mech;Ranjitha Kumar
通讯作者: Thomas F. Liu;Mark Craft;Jason Situ;Ersin Yumer;R. Mech;Ranjitha Kumar
PUMICE:从自然语言和演示中学习概念和条件的多模式代理
DOI: 10.1145/3332165.3347899
发表时间: 2019
期刊: UIST'19
影响因子: --
作者:
Li, Toby Jia-Jun;Radensky, Marissa;Jia, Justin;Singarajah, Kirielle;Mitchell, Tom M.;Myers, Brad A.
通讯作者: Myers, Brad A.