A Framework and Tool for Collaborative Extraction of Reliable Information

A Framework and Tool for Collaborative Extraction of Reliable Information
复制标题

DOI:
--
复制
发表时间:
2013-10
期刊:
--
影响因子:
--
通讯作者:
Graham Neubig;Shinsuke Mori;M. Mizukami
Graham Neubig;Shinsuke Mori;M. Mizukami
中科院分区:
其他
文献类型:
--
作者:
Graham Neubig;Shinsuke Mori;M. Mizukami

文献摘要

相似文献

这项研究提出了一个有效的信息提取和过滤的框架,在这种情况下,1)极端的可靠性是重要的,2)要梳理的信息量是巨大的,3)我们可以预期有相对大量的人类工作人员可用。特别是,我们受到危机时期需求的激励,并假设为了确保所需的高可靠性,至少需要一名人类工作人员确认所有提取的信息。在这种情况下,我们提出了一种方法来提高人工验证的效率,通过使用机器学习技术来决定向工人提供哪些信息。即使有了这个高效的搜索框架,互联网上的信息量仍然太多,一个用户无法处理,所以我们额外创建了一个基于Web的框架,允许协作工作,以及一个算法,允许这个框架实时处理大数据。我们使用东日本大地震后的Twitter数据进行评估,并比较使用传统关键字搜索和基于学习的方法的效率。
This research proposes a framework for efficient information extraction and filtering in situations where 1) extreme reliability is important, 2) the amount of information to be combed through is massive, and 3) we can expect a relatively large number of human workers to be available. In particular, we are motivated by needs in times of crisis, and assume that in order to ensure the high level of reliability required, it will be necessary to have at least one human worker confirm all extracted information. Given this setting, we propose a method to improve the efficiency of manual verification by deciding which information to present to workers using machine learning techniques. Even given this efficient search framework, the amount of information on the internet is still too much for one user to handle, so we additionally create a web-based framework that allows for collaborative work, and an algorithm that allows for this framework to work on large data in real-time. We perform an evaluation using data from Twitter after the Great East Japan Earthquake, and compare efficiency using both traditional keyword search and the proposed learningbased method.