CAREER: Developing an Underspecified Representation for Temporal Information in Text
CAREER: Developing an Underspecified Representation for Temporal Information in Text
批准号:
1652742
负责人:
Anna Rumshisky
金额:
$49.94万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-03-15 至 2024-02-29
中文摘要
尽管最近在自然语言的自动处理方面取得了进展,但在许多情况下,接近一般人类水平的文本理解仍然具有挑战性。这个CAREER项目解决了这样一个挑战,自动提取和理解自然语言文本叙述中描述的事件的顺序和时间。它开发了一个用于表示和提取文本中传达的时间信息的计算框架,其最终目标是实现从文本中进行现实的时间推理。它还鼓励学生尽早参与计算机科学研究,特别是旨在吸引女学生从事计算机科学和相关领域的职业。与现有方法不同的是,该研究假设规格不足是时间表征的一个整体属性,并由粗粒度事件聚类的概念提供支持。它利用了叙事文本的微观结构,通过识别事件集群及其叙事锚点,以及默认的事件时间和持续时间,来组织时间线的未明确表示。它还通过开发三个关键组成部分来解决当前最先进技术中的知识差距。首先,采用了一种新颖的表征方案,为注释者提供了简单、直观的选择,从而最大限度地减少了认知努力并减少了注释错误。问答作为目标应用程序,重点关注阅读理解问题,这些问题需要超越理解简单事实的时间推理。其次,使用新的方法对标注一致性进行内在评价,以解决现有评价方法存在的问题,这些方法通常会根据对时间关系图通过传递闭包推断出的时间关系的具体评价策略而产生不同的结果。所提出的方法依赖于将时间关系表示为偏序集,并使用这些偏序的线性扩展来评估注释器间的一致性和系统性能。最后,研究了一类新的基于神经网络的模型,旨在恢复锚定事件聚类上的偏序图。这些模型使用外部记忆组件进行累积话语表示,并允许联合训练来识别共同指代的事件提及,将恢复的事件分组为大致同时的事件集群,并在它们之间建立类型链接。这些模型包含默认事件顺序和时间的表示,以及事件的参数结构和时间表达式的定量推理。
英文摘要
Despite the recent advances in automated processing of natural language, approaching general human-level understanding of text in many cases still remains challenging. This CAREER project addresses one such challenge, automatically extracting and understanding the order and timing of events described in natural language text narratives. It develops a computational framework for representing and extracting temporal information conveyed in text, with the end goal to enable realistic temporal reasoning from text. It also engages student involvement in computer science research early on, and in particular is designed to attract female students to pursue careers in computer science and related areas.As distinct from existing approaches, the proposed research assumes underspecification to be an integral property of temporal representation, supported by the notion of a coarse-grained event cluster. It takes advantage of the micro-structure of narrative text by identifying event clusters and their narrative anchors which, together with default event times and durations, serve to organize the underspecified representation of the timeline. It also addresses knowledge gaps in the current state-of-the-art in by developing three key components. First, a novel representational scheme is employed to facilitate simple, intuitive choices for the annotators that minimize cognitive effort and reduce annotation error. Question answering serves as the target application, with a focus on reading-comprehension questions that require temporal reasoning beyond understanding simple factoids. Second, novel methods for intrinsic evaluation of annotation consistency are used to address problems with existing evaluation methods, which often produce varied results depending on specific strategies for crediting temporal relations inferred through transitive closure over the temporal relation graph. The proposed methods rely on representing temporal relations as partially ordered sets and use linear extensions of these partial orders in order to evaluate inter-annotator agreement and system performance. Finally, a new class of neural network-based models is explored that aim to recover partial order graphs over anchored event clusters. These models use external memory components for cumulative discourse representation, and allow joint training for identifying coreferent event mentions, grouping the recovered events into roughly-simultaneous event clusters, and establishing typed links between them. The models incorporate a representation for default event order and timing, as well as argument structure for events and quantitative inference over temporal expressions.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.18653/v1/p18-1049
发表时间:
2018-07
期刊:
影响因子:
--
作者:
[Yuanliang Meng;Anna Rumshisky]
通讯作者:
Yuanliang Meng;Anna Rumshisky
DOI:
10.1609/aaai.v34i05.6398
发表时间:
2020-04
期刊:
影响因子:
--
作者:
[Anna Rogers;Olga Kovaleva;Matthew Downey;Anna Rumshisky]
通讯作者:
Anna Rogers;Olga Kovaleva;Matthew Downey;Anna Rumshisky
Down and Across: Introducing Crossword-Solving as a New NLP Benchmark
纵横交错:引入填字游戏作为新的 NLP 基准
DOI:
10.18653/v1/2022.acl-long.189
发表时间:
2022
期刊:
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers
影响因子:
--
作者:
[Kulshreshtha, Saurabh, Kovaleva, Olga, Shivagunde, Namrata, Rumshisky, Anna]
通讯作者:
Rumshisky, Anna
DOI:
--
发表时间:
2018-08
期刊:
ArXiv
影响因子:
--
作者:
[Yuanliang Meng;Anna Rumshisky]
通讯作者:
Yuanliang Meng;Anna Rumshisky
DOI:
10.18653/v1/d17-1092
发表时间:
2017-03
期刊:
ArXiv
影响因子:
--
作者:
[Yuanliang Meng;Anna Rumshisky;Alexey Romanov]
通讯作者:
Yuanliang Meng;Anna Rumshisky;Alexey Romanov
Collaborative Research: Machine Learning for Student Reasoning during Challenging Concept Questions
-
批准号:2226601
-
项目类别:Standard Grant
-
资助金额:$17.17万
-
财政年份:2023
-
负责人:Anna Rumshisky
-
依托单位:
Student Participant Support for Conversational Intelligence Summer School 2019
-
批准号:1933903
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:2019
-
负责人:Anna Rumshisky
-
依托单位:
EAGER: Exploring Cognitively Plausible Computational Models for Processing Human Language
-
批准号:1844740
-
项目类别:Standard Grant
-
资助金额:$10.99万
-
财政年份:2018
-
负责人:Anna Rumshisky
-
依托单位:
海外基金