课题基金 / 基金详情

CAREER: Multilingual Learning for Event Structures from Text

CAREER: Multilingual Learning for Event Structures from Text
职业:从文本中学习事件结构的多语言
批准号:
2239570
负责人:
Thien Nguyen
金额:
$58.22万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-06-01 至 2028-05-31

项目摘要

项目成果

Thien Nguyen的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Natural language text is replete with important events in different areas (protests, cybersecurity breaches, elections, disease outbreaks, and business transactions). Identifying events to describe who did what to whom and their relations (causal, subevent, and coreferential) from a large amount of text can provide valuable data to support intelligent applications and data-driven decisions over various domains. However, current event structure extraction systems can only perform over text data for a few popular languages such as English, Chinese, Spanish, and Arabic. Text data from many other languages in the world thus cannot be processed by current event extraction systems. This limitation has hindered the coverage of data sources for the systems, introduced language biases in the extracted events, and delayed updates with latest events in local reports. Eventually, the collected event data from current techniques cannot comprehensively represent the latest dynamics over the world to effectively support decision making for important problems of national interests. To address the multilingual challenges, this project will develop event extraction and event-event relation extraction systems that can be effective for data in multiple languages, emphasizing on understudied and low-resource languages to improve the coverage of extracted data and promote democratization of technologies. In information retrieval, multilingual event structure data from the developed technologies can enable data management systems to quickly obtain answers and create summaries for broader user queries in many more languages. In cybersecurity, databases for extracted cyber attack events from multilingual sources can be used to generate more fine-grained and comprehensive reports to inform resource allocation decisions to better protect online activities. In socio-political science, coded conflict and meditation events from more languages can increase the scope and reduce biases of the data to support better decisions for foreign policy, civil war prevention, environmental challenges, or economic strategies.This project will address three fundamental limitations of existing multilingual learning research for event structure extraction: (i) the lack of multilingual datasets that provide data annotation for multiple languages to sufficiently support generalization evaluation of models across different language families, (ii) the limitations of current multilingual representation learning methods when aligning representations between languages to induce language-general features, and (iii) the scarcity of labeled data in different languages to train multilingual models. First, the project will annotate documents for all event extraction and event-event relation extraction tasks in many more languages using consistent schemas. The selected languages for annotation will be typologically diverse, understudied and low-resource to provide reliable multilingual evaluation data for the developed methods. Second, to boost cross-lingual performance for event structure extraction, this project will devise multilingual representation learning methods to enable effective knowledge transfer where models trained on labeled data of high-resource languages can be directly applied to data of other languages. The project will develop novel representation alignment methods for different languages using representation matching, augmentation, and language-general structure induction for text. Third, concerning limited training data for multilingual learning, this project will develop novel methods to automatically generate labeled data in different languages. The project will introduce techniques to mitigate noises in the generated data and optimize generation procedures to boost multilingual learning and performance. The research activities in this project will be closely integrated with education and outreach missions to broaden their impacts.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
DOI: 10.18653/v1/2023.findings-acl.727
发表时间: 2023
期刊:
影响因子: --
作者: [Amir Pouran Ben Veyseh;Franck Dernoncourt;Bonan Min;Thien Huu Nguyen]
通讯作者: Amir Pouran Ben Veyseh;Franck Dernoncourt;Bonan Min;Thien Huu Nguyen
DOI: 10.18653/v1/2023.findings-acl.135
发表时间: 2023
期刊:
影响因子: --
作者: [Chien Nguyen;Linh Van Ngo;Thien Huu Nguyen]
通讯作者: Chien Nguyen;Linh Van Ngo;Thien Huu Nguyen
DOI: 10.18653/v1/2023.acl-long.296
发表时间: 2023
期刊:
影响因子: --
作者: [Luis Guzman Nateras;Franck Dernoncourt;Thien Huu Nguyen]
通讯作者: Luis Guzman Nateras;Franck Dernoncourt;Thien Huu Nguyen
DOI: 10.18653/v1/2023.findings-acl.266
发表时间: 2023
期刊:
影响因子: --
作者: [Chien Van Nguyen;Hieu Man;Thien Huu Nguyen]
通讯作者: Chien Van Nguyen;Hieu Man;Thien Huu Nguyen
Phase I IUCRC University of Oregon: Center for Big Learning
  • 批准号:
    1747798
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $75.0万
  • 财政年份:
    2018
  • 负责人:
    Thien Nguyen
  • 依托单位:
海外基金