Content and Context: Two-Pronged Bootstrapped Learning for Regex-Formatted Entity Extraction

Content and Context: Two-Pronged Bootstrapped Learning for Regex-Formatted Entity Extraction
复制标题

内容和上下文:正则表达式格式实体提取的双管齐下引导学习

DOI:
--
复制
发表时间:
2018
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
S. Mehta
S. Mehta
中科院分区:
--
文献类型:
--
作者:
Stanley Simoes;P Deepak;Munu Sairamesh;D. Khemani;S. Mehta

文献摘要

被引文献

相似文献

正则表达式是基于规则的信息提取系统的重要构建块。正则表达式可以对规则进行编码以识别简单实体的实例,然后将其输入到更复杂的跨实体关系的识别中。手动制作一个识别实体所有可能实例的正则表达式很困难,因为实体可以以各种不同的形式出现。因此,自动泛化手动制作的种子正则表达式以提高 IE 系统的召回率的问题引起了研究关注。在本文中,我们提出了一种引导方法来提高正则表达式格式实体提取的召回率,唯一的监督来源是种子正则表达式。我们的方法从为感兴趣的实体手动编写的高精度种子正则表达式开始,并使用种子正则表达式的匹配以及这些匹配周围的上下文来识别该实体的更多实例。然后使用它们来识别代表该实体的一组多样化、高召回率的正则表达式。通过对多个现实世界文档语料库的实证评估,我们说明了我们方法的有效性。
Regular expressions are an important building block of rule-based information extraction systems. Regexes can encode rules to recognize instances of simple entities which can then feed into the identification of more complex cross-entity relationships. Manually crafting a regex that recognizes all possible instances of an entity is difficult since an entity can manifest in a variety of different forms. Thus, the problem of automatically generalizing manually crafted seed regexes to improve the recall of IE systems has attracted research attention. In this paper, we propose a bootstrapped approach to improve the recall for extraction of regex-formatted entities, with the only source of supervision being the seed regex. Our approach starts from a manually authored high precision seed regex for the entity of interest, and uses the matches of the seed regex and the context around these matches to identify more instances of the entity. These are then used to identify a set of diverse, high recall regexes that are representative of this entity. Through an empirical evaluation over multiple real world document corpora, we illustrate the effectiveness of our approach.