Precise temporal slot filling via truth finding with data-driven commonsense

Precise temporal slot filling via truth finding with data-driven commonsense
复制标题

通过数据驱动的常识发现真相来精确填充时隙

DOI:
10.1007/s10115-020-01493-w
复制
发表时间:
2020
影响因子:
2.7
通讯作者:
Jiang, Meng
Jiang, Meng
中科院分区:
计算机科学4区
文献类型:
--
作者:
Wang, Xueying;Jiang, Meng

文献摘要

参考文献

相似文献

时态槽填充(TSF)的任务是从文本数据中提取给定实体的特定属性值(称为“事实”)以及事实的时态标签。现有的工作将时间标签表示为单个时隙,本文介绍并研究了精确TSF(PTSF)的任务,即填充包括开始和结束时间点的两个精确时隙。根据我们对新闻语料库的观察,大部分的事实都应该有这两点,但只有不到0.1%的事实有时间表达。另一方面,文件的发布时间虽然经常可用,但并不像事实有效时间的时间表达那样精确。因此,直接分解时间表达式或使用任意的后时间段不能为PTSF提供准确的结果。PTSF的挑战在于在文本中嘈杂和不完整的时间背景中找到精确的时间标签。为了应对这一挑战,我们提出了一种基于真理发现哲学的无监督方法。该方法有两个模块,相互增强:一个是有条件的时间上下文的事实提取器的可靠性估计;另一个是基于提取器的可靠性的事实可信度估计。常识知识(例如,一个国家在特定时间只有一位总统)是从数据中自动生成的,并用于根据可信的事实推断虚假的主张。为了评估的目的,我们从维基百科手动收集了数百个时间事实作为基础事实,包括国家的总统任期和运动队的球员职业生涯历史。在一个大型新闻数据集上的实验证明了该算法的准确性和有效性。
The task of temporal slot filling (TSF) is to extract values of specific attributes for a given entity, called “facts”, as well as temporal tags of the facts, from text data. While existing work denoted the temporal tags as single time slots, in this paper, we introduce and study the task of Precise TSF (PTSF), that is to fill two precise temporal slots including the beginning and ending time points. Based on our observation from a news corpus, most of the facts should have the two points, however, fewer than 0.1% of them have time expressions in the documents. On the other hand, the documents’ post time, though often available, is not as precise as the time expressions of being the time a fact was valid. Therefore, directly decomposing the time expressions or using an arbitrary post-time period cannot provide accurate results for PTSF. The challenge of PTSF lies in finding precise time tags in noisy and incomplete temporal contexts in the text. To address the challenge, we propose an unsupervised approach based on the philosophy of truth finding. The approach has two modules that mutually enhance each other: One is a reliability estimator of fact extractors conditionally on the temporal contexts; the other is a fact trustworthiness estimator based on the extractor’s reliability. Commonsense knowledge (e.g., one country has only one president at a specific time) was automatically generated from data and used for inferring false claims based on trustworthy facts. For the purpose of evaluation, we manually collect hundreds of temporal facts from Wikipedia as ground truth, including country’s presidential terms and sport team’s player career history. Experiments on a large news dataset demonstrate the accuracy and efficiency of our proposed algorithm.
DOI: 10.1145/3219819.3220017
发表时间: 2018-07
期刊: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
影响因子: --
作者:
Qi Li;Meng Jiang;Xikun Zhang;Meng Qu;T. Hanratty;Jing Gao;Jiawei Han
通讯作者: Qi Li;Meng Jiang;Xikun Zhang;Meng Qu;T. Hanratty;Jing Gao;Jiawei Han
DOI: 10.1145/2983323.2983751
发表时间: 2016-10
期刊: Proceedings of the 25th ACM International on Conference on Information and Knowledge Management
影响因子: --
作者:
Tuan-Anh Hoang-Vu;H. Vo;J. Freire
通讯作者: Tuan-Anh Hoang-Vu;H. Vo;J. Freire
使用集成真理发现方法估计数据准确性
DOI: --
发表时间: 2015
期刊: 2015 IEEE International Conference on Big Data (Big Data)
影响因子: --
作者:
Laure Berti
通讯作者: Laure Berti
DOI: 10.1145/3132847.3133038
发表时间: 2017-11
期刊: Proceedings of the 2017 ACM on Conference on Information and Knowledge Management
影响因子: --
作者:
M. Chekol
通讯作者: M. Chekol
DOI: 10.1145/3097983.3098105
发表时间: 2017-03
期刊: Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
影响因子: --
作者:
Meng Jiang;Jingbo Shang;Taylor Cassidy;Xiang Ren;Lance M. Kaplan;T. Hanratty;Jiawei Han
通讯作者: Meng Jiang;Jingbo Shang;Taylor Cassidy;Xiang Ren;Lance M. Kaplan;T. Hanratty;Jiawei Han