TruePIE: Discovering Reliable Patterns in Pattern-Based Information Extraction

TruePIE: Discovering Reliable Patterns in Pattern-Based Information Extraction
复制标题

DOI:
10.1145/3219819.3220017
复制
发表时间:
2018-07
期刊:
Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
Qi Li;Meng Jiang;Xikun Zhang;Meng Qu;T. Hanratty;Jing Gao;Jiawei Han
Qi Li;Meng Jiang;Xikun Zhang;Meng Qu;T. Hanratty;Jing Gao;Jiawei Han
中科院分区:
其他
文献类型:
--
作者:
Qi Li;Meng Jiang;Xikun Zhang;Meng Qu;T. Hanratty;Jing Gao;Jiawei Han

文献摘要

被引文献

相似文献

基于模式的方法在信息抽取和自然语言处理研究中取得了成功。先前的方法基于文本模式的个体内容的统计来学习文本模式的质量作为与特定任务的相关性(例如,长度,频率)和数百个仔细注释的标签。然而,由于相关性和正确性之间的巨大差距,良好的内容质量模式可能会产生严重的信息冲突。在(实体、属性、值)元组抽取中,信息的正确性评价是一个关键问题.在这项工作中,我们提出了一种新的方法,称为TruePIE,找到可靠的模式,不仅可以提取相关的,但也正确的信息。TruePIE采用自训练框架,重复训练-预测-提取过程,逐渐发现越来越多的可靠模式。为了更好地表示文本模式,模式嵌入公式化,使得具有相似语义的模式彼此紧密地嵌入。嵌入同时考虑了提取的局部模式信息和分布信息。为了克服缺乏对模式可靠性的监督的挑战,TruePIE可以通过应用arity约束来区分高可靠模式(即,正模式)和高度不可靠的模式(即,负模式)。在一个巨大的新闻数据集(超过25 GB)上的实验表明,所提出的TruePIE在三个任务中的每一个上都显着优于基线方法:可靠的元组提取,可靠的模式提取和否定模式提取。
Pattern-based methods have been successful in information extraction and NLP research. Previous approaches learn the quality of a textual pattern as relatedness to a certain task based on statistics of its individual content (e.g., length, frequency) and hundreds of carefully-annotated labels. However, patterns of good content-quality may generate heavily conflicting information due to the big gap between relatedness and correctness. Evaluating the correctness of information is critical in (entity, attribute, value)-tuple extraction. In this work, we propose a novel method, called TruePIE, that finds reliable patterns which can extract not only related but also correct information. TruePIE adopts the self-training framework and repeats the training-predicting-extracting process to gradually discover more and more reliable patterns. To better represent the textual patterns, pattern embeddings are formulated so that patterns with similar semantic meanings are embedded closely to each other. The embeddings jointly consider the local pattern information and the distributional information of the extractions. To conquer the challenge of lacking supervision on patterns' reliability, TruePIE can automatically generate high quality training patterns based on a couple of seed patterns by applying the arity-constraints to distinguish highly reliable patterns (i.e., positive patterns) and highly unreliable patterns (i.e., negative patterns). Experiments on a huge news dataset (over 25GB) demonstrate that the proposed TruePIE significantly outperforms baseline methods on each of the three tasks: reliable tuple extraction, reliable pattern extraction, and negative pattern extraction.