课题基金 / 基金详情

EAGER: Building Idiomaticity into Natural Language Processing

EAGER: Building Idiomaticity into Natural Language Processing
EAGER:将惯用性融入自然语言处理
批准号:
2230817
负责人:
Suma Bhat
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
已结题
起止时间:
2022-08-15 至 2024-07-31

项目摘要

项目成果

Suma Bhat的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Idiomatic expressions are an essential component of everyday language use and the hallmark of native language ability. Consider the phrase throw away; proficient speakers can effortlessly understand that the phrase takes a figurative meaning in “Britain threw away all the achievements of the last decade.” and a literal sense in “He threw away his cigarette and buried his head in his arms.” This EArly Grant for Exploratory Research (EAGER) will build a high-quality dataset for computers to understand the differences between figurative and literal senses of these expressions in general English text. The main novelty of this project will be in collecting a large class of idiomatic expressions and sentences containing them to let computers learn the inherent variability between a variety of idiomatic phrases. Collecting many sentences with phrases that have a figurative and literal meaning will permit computers better understand the nuances with which these expressions are used in everyday conversations and writing. Beyond understanding them, the collected. examples will help computers use these expressions like native speakers do when automatically writing text and even suggest appropriate expressions in specific contexts.This EAGER project is essentially interdisciplinary spanning the areas of linguistics and computation and will investigate novel paradigms for natural language processing that are idiomaticity-aware. As such, it will have two research aims: (1) creating a high-quality dataset of phrasal verbs annotated with their context-specific senses and their literal/figurative equivalent forms, and (2) testing the performance of state-of-the-art idiomaticity-aware algorithms. Because idiomatic expressions vary widely in form and structure, the focus on phrasal verbs (also known as verb-particle constructions) in the context of the exploratory project will permit studying a very frequent class of idiomatic expressions that are syntactically different from those in currently available datasets. The primary risk of this project stems from its exploratory nature of creating large corpora with sufficient coverage for language model training. Given their prevalence in natural language, the dataset of phrasal verbs in English will supplement available datasets on idiomatic expressions in terms of their variety. Moreover, their figurative and literal ambiguity in context (apart from their polysemy) will permit a diverse look at the phenomenon of non-compositionality that characterizes idiomatic expressions. Thus, the dataset will serve as a training and test bed for algorithms that detect, interpret, and generate a broad class of idiomatic expressions. This effort will lead to new natural language processing algorithms for accurate interpretation and generation of idiomatic expressions towards a more human-like language processing ability in machines.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: Collaborative: BystanderBots: Automated Bystander Intervention for Cyberbullying Mitigation
国内基金
海外基金
基于支链淀粉building blocks构建优质BE突变酶定向修饰淀粉调控机制的研究
  • 批准号:
    31771933
  • 项目类别:
    面上项目
  • 资助金额:
    60.0万元
  • 批准年份:
    2017
  • 负责人:
    郭丽
  • 依托单位: