Schema2QA: High-Quality and Low-Cost Q&A Agents for the Structured Web

Schema2QA: High-Quality and Low-Cost Q&A Agents for the Structured Web
复制标题

Schema2QA:高质量、低成本的 Q

DOI:
10.1145/3340531.3411974
复制
发表时间:
2020
期刊:
CIKM '20: The 29th ACM International Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Lam, Monica S.
Lam, Monica S.
中科院分区:
--
文献类型:
--
作者:
Xu, Silei;Campagna, Giovanni;Li, Jian;Lam, Monica S.

文献摘要

参考文献

被引文献

相似文献

构建问答代理目前需要大量的带注释的数据集,而这些数据集的成本高得令人望而却步。本文提出了一个开源的工具包--Schema2QA,它可以从数据库模式中生成一个问答系统,并为每个字段添加一些注释。其关键概念是通过在通用查询模板语料库的帮助下合成大量领域内问题来覆盖数据库上可能的复合查询空间。利用人工合成的数据和一个较小的释义集来训练一个基于BERT预训练模型的新型神经网络。我们使用Schema2QA对5个Schema.org域、餐馆、人物、电影、书籍和音乐进行了问答系统的生成,在这些域的众包问题上获得了%到75%的总体准确率。一旦获得了模式的注释和释义,就不需要额外的手动工作来为使用相同模式的任何网站创建Q&A代理。此外,我们证明了学习可以从餐馆领域转移到酒店领域,不需要人工努力就可以在众包问题上获得%的准确率。在可以使用Schema.org回答的热门餐厅问题上,Schema2QA的准确率达到了60%。它的性能与谷歌助手不相上下,比Siri低7%,比Alexa高15%。在更复杂、更长尾的问题上,它的表现比所有这些助手至少高出18%。
Building a question-answering agent currently requires large annotated datasets, which are prohibitively expensive. This paper proposes Schema2QA, an open-source toolkit that can generate a Q&A system from a database schema augmented with a few annotations for each field. The key concept is to cover the space of possible compound queries on the database with a large number of in-domain questions synthesized with the help of a corpus of generic query templates. The synthesized data and a small paraphrase set are used to train a novel neural network based on the BERT pretrained model. We use Schema2QA to generate Q&A systems for five Schema.org domains, restaurants, people, movies, books and music, and obtain an overall accuracy between 64% and 75% on crowdsourced questions for these domains. Once annotations and paraphrases are obtained for a Schema.org schema, no additional manual effort is needed to create a Q&A agent for any website that uses the same schema. Furthermore, we demonstrate that learning can be transferred from the restaurant to the hotel domain, obtaining a 64% accuracy on crowdsourced questions with no manual effort. Schema2QA achieves an accuracy of 60% on popular restaurant questions that can be answered using Schema.org. Its performance is comparable to Google Assistant, 7% lower than Siri, and 15% higher than Alexa. It outperforms all these assistants by at least 18% on more complex, long-tail questions.
DOI: 10.1145/3318464.3380589
发表时间: 2020-05
期刊: Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data
影响因子: --
作者:
Nathaniel Weir;Prasetya Ajie Utama;Alex Galakatos;Andrew Crotty;Amir Ilkhechi;Shekar Ramaswamy;Rohin Bhushan;Nadja Geisler;Benjamin Hättasch;Steffen Eger;U. Çetintemel;Carsten Binnig
通讯作者: Nathaniel Weir;Prasetya Ajie Utama;Alex Galakatos;Andrew Crotty;Amir Ilkhechi;Shekar Ramaswamy;Rohin Bhushan;Nadja Geisler;Benjamin Hättasch;Steffen Eger;U. Çetintemel;Carsten Binnig
DOI: 10.18653/v1/2020.acl-main.12
发表时间: 2020-05
期刊: ArXiv
影响因子: --
作者:
Giovanni Campagna;Agata Foryciarz;M. Moradshahi;M. Lam
通讯作者: Giovanni Campagna;Agata Foryciarz;M. Moradshahi;M. Lam
从树库中引入确定性 Prolog 解析器:一种机器学习方法
DOI: --
发表时间: 1994
期刊: AAAI Conference on Artificial Intelligence
影响因子: --
作者:
J. Zelle;R. Mooney
通讯作者: R. Mooney