CAREER: Drawing inferences for human-like language understanding
CAREER: Drawing inferences for human-like language understanding
批准号:
1845122
负责人:
Marie-Catherine de Marneffe
金额:
$49.99万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-08-15 至 2024-07-31
中文摘要
在处理语言时,读者和听众不仅仅理解他们读到或听到的单词的字面意思;他们还从中得出推论。例如,如果有人告诉你:“我说你在这个时候过来是疯了。这是一个世界性的事件。你知道威尼斯到处都是游客吗?“,他们可能会推断威尼斯确实挤满了游客。但是在“她这样多久了?你看医生了吗?你知不知道这病是治不好的?”,他们不会推断这是不治之症,即使这两个事件都在一个问题中,并嵌入在同一串单词“你知道吗”下。不同的因素,如使用的语气或世界知识,在得出这些推论时发挥作用。该项目旨在研究这些因素,并开发自动获取推论的广泛覆盖模型。这些模型对需要准确推理过程的自然语言处理(NLP)任务(如信息提取)具有影响。此外,为了实现类似人类的语言理解,NLP技术不仅要开发模型来捕捉语言中传达的内容,而不是明确地说出来,而且还要评估推理对大多数人来说是否是系统的,或者是否会出现不同的解释。该项目研究了如何将来自人们直觉的“常识”数据中存在的可变性准确地表示在当前构建NLP系统的数据集中,从而捕获这些数据。最近,NLP领域的大量工作都集中在深度学习、新任务和基准测试上。然而,这样的冒险无助于理解人类语言的细节,也无助于确定哪些特征对语言处理真正重要。该项目针对不同集合中的分类和非分类推理:关于情感,协议和扬声器承诺(扬声器是否致力于他们所描述的事件的真相)的推理,并重新定义实现类人自然语言理解所需的基准。它研究了数据驱动方法和使用专业语言特征之间的更好协同作用如何导致NLP系统的根本性进步。该项目还定量研究了大量自然发生的例子上语言特征的相互作用,因此不仅对NLP而且对语言理论都有影响。结果将包括更好地掌握如何使用语言见解来自动实现人类水平的理解;公开可用的数据比当前数据集更适合人类对语言的直觉,从而可以用于锐化NLP模型;为学生提供的课程材料和为公众提供的演示,以提高对社交媒体所产生的社会问题的认识,强调在日常交流中,语言所传达的信息的重要性,而不仅仅是一串明确的单词,该奖项反映了NSF的法定使命,并通过使用基金会的知识产权进行评估,被认为值得支持。优点和更广泛的影响审查标准。
英文摘要
When dealing with language, readers and listeners understand more than just the literal meaning of the words they read or hear; they also draw inferences from them. For instance, if someone tells you "I said you were mad to come over at this time. It's a world event. Do you know that Venice is packed with visitors?", they will likely infer that Venice is indeed packed with visitors. However in "How long has she been like this? Did you see a doctor? Do you know that it is incurable?", they will not infer that it is incurable, even though both events are in a question and embedded under the same string of words "do you know". Different factors, like the tone used or world knowledge, play a role in deriving these inferences. The project aims at studying these factors and developing broad-coverage models that automatically capture inferences. Such models have implications for natural language processing (NLP) tasks that require an accurate inference process, such as information extraction. Further, to achieve human-like language understanding, it is not only crucial for NLP technologies to develop models that capture what is conveyed in language without being explicitly said, but to also assess whether the inferences are systematic for most people, or whether different interpretations arise. This project investigates how the variability present in "common sense" data that come from people's intuitions can be accurately represented in the type of datasets on which NLP systems are currently built, and thereby be captured. Recently a large body of work in NLP has focused on deep learning, hill-climbing on new tasks and benchmarks. However such ventures do not help with the understanding of the details of human language or in determining which features actually matter for language processing. This project targets both categorical and non-categorical inferences in a diverse set: inferences about sentiment, agreement and speaker commitment (whether speakers are committed to the truth of the events they describe), and redefining the kind of benchmarks needed to achieve human-like natural language understanding. It investigates how a better synergy between data-driven methods and the use of specialized linguistic features can lead to fundamental advances in NLP systems. The project also quantitatively studies the interactions of linguistic features on a large amount of naturally occurring examples, and has thus an impact not only for NLP but also for linguistic theories. Results will include a better grasp of how linguistic insights can be used to automatically achieve human-level understanding; a publicly available data that better fit human intuitions on language than current datasets do and can thereby serve to sharpen NLP models; and course materials for students and demos for the general public that raise awareness of societal problems engendered by social media, emphasize the importance of what gets conveyed by language beyond the explicit string of words in everyday communication, and demonstrate what can be achieved when research in linguistics and computer science is combined.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.18148/sub/2020.v24i2.884
发表时间:
2020
期刊:
Proceedings of Sinn und Bedeutung 24
影响因子:
--
作者:
[Mahler, Taylor, de Marneffe, Marie-Catherine, Lai, Catherine]
通讯作者:
Lai, Catherine
DOI:
10.1162/tacl_a_00523
发表时间:
2022-09
期刊:
Transactions of the Association for Computational Linguistics
影响因子:
10.9
作者:
[Nan Jiang;M. Marneffe]
通讯作者:
Nan Jiang;M. Marneffe
He Thinks He Knows Better than the Doctors: BERT for Event Factuality Fails on Pragmatics
他认为他比医生更了解:事件事实性的 BERT 在语用学上失败
DOI:
10.1162/tacl_a_00414
发表时间:
2021
期刊:
Transactions of the Association for Computational Linguistics
影响因子:
10.9
作者:
[Jiang, Nanjiang, de Marneffe, Marie-Catherine]
通讯作者:
de Marneffe, Marie-Catherine
DOI:
10.18653/v1/2021.naacl-main.390
发表时间:
2021
期刊:
Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
影响因子:
--
作者:
[Zhang, Xinliang Frederick, de Marneffe, Marie-Catherine]
通讯作者:
de Marneffe, Marie-Catherine
Evaluating BERT for natural language inference: A case study on the CommitmentBank
评估 BERT 的自然语言推理能力:CommitmentBank 的案例研究
DOI:
10.18653/v1/d19-1630
发表时间:
2019
期刊:
Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP
影响因子:
--
作者:
[Jiang, Nanjiang, de Marneffe, Marie-Catherine]
通讯作者:
de Marneffe, Marie-Catherine
Student travel support to the Fourth Universal Dependencies Workshop (2020)
-
批准号:2024161
-
项目类别:Standard Grant
-
资助金额:$0.6万
-
财政年份:2020
-
负责人:Marie-Catherine de Marneffe
-
依托单位:
2018 Association for Computational Linguistics (ACL) Student Workshop
-
批准号:1827830
-
项目类别:Standard Grant
-
资助金额:$1.8万
-
财政年份:2018
-
负责人:Marie-Catherine de Marneffe
-
依托单位:
CRII: RI: What do you mean? -- Automatic identification of inferences drawn from text
-
批准号:1464252
-
项目类别:Standard Grant
-
资助金额:$14.34万
-
财政年份:2015
-
负责人:Marie-Catherine de Marneffe
-
依托单位:
Collaborative Research: What's the question? A cross-linguistic investigation into compositional and pragmatic constraints on the question under discussion
-
批准号:1452674
-
项目类别:Standard Grant
-
资助金额:$27.31万
-
财政年份:2015
-
负责人:Marie-Catherine de Marneffe
-
依托单位:
海外基金