CAREER: Multilingual Learning for Event Structures from Text
CAREER: Multilingual Learning for Event Structures from Text
批准号:
2239570
负责人:
Thien Nguyen
金额:
$58.22万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-06-01 至 2028-05-31
中文摘要
自然语言文本充满了不同领域的重要事件(抗议,网络安全漏洞,选举,疾病爆发和商业交易)。从大量文本中识别事件来描述谁对谁做了什么以及它们的关系(因果关系,子事件和共指关系)可以提供有价值的数据来支持智能应用程序和各种领域的数据驱动决策。然而,当前的事件结构提取系统只能对少数流行语言(诸如英语、中文、西班牙语和阿拉伯语)的文本数据执行。来自世界上许多其他语言的文本数据因此不能被当前的事件提取系统处理。这一限制阻碍了系统数据源的覆盖,在提取的事件中引入了语言偏见,并延迟了本地报告中最新事件的更新。最终,从现有技术中收集的事件数据不能全面代表世界各地的最新动态,以有效地支持国家利益的重大问题的决策。为了应对多语言的挑战,本项目将开发事件提取和事件-事件关系提取系统,这些系统可以有效地处理多种语言的数据,重点是研究不足和资源不足的语言,以提高提取数据的覆盖率并促进技术的民主化。在信息检索中,来自已开发技术的多语言事件结构数据可以使数据管理系统能够快速获得答案,并以更多语言为更广泛的用户查询创建摘要。在网络安全方面,从多语言来源提取的网络攻击事件数据库可用于生成更细粒度和更全面的报告,为资源分配决策提供信息,以更好地保护在线活动。在社会政治科学中,来自更多语言的编码冲突和冥想事件可以增加数据的范围并减少数据的偏差,以支持外交政策,内战预防,环境挑战或经济战略的更好决策。该项目将解决现有多语言学习研究的三个基本局限性,用于事件结构提取:(i)缺乏为多种语言提供数据注释的多语言数据集,以充分支持跨不同语系的模型的泛化评估,(ii)当前多语言表示学习方法在对齐语言之间的表示以诱导语言通用特征时的局限性,以及(iii)缺乏不同语言的标记数据来训练多语言模型。首先,该项目将使用一致的模式为更多语言的所有事件提取和事件-事件关系提取任务注释文档。选定的注释语言将在类型上多样化,研究不足,资源少,为开发的方法提供可靠的多语言评估数据。其次,为了提高事件结构提取的跨语言性能,该项目将设计多语言表示学习方法,以实现有效的知识转移,其中在高资源语言的标记数据上训练的模型可以直接应用于其他语言的数据。该项目将使用文本的表示匹配、增强和语言通用结构归纳,为不同的语言开发新的表示对齐方法。第三,针对多语言学习的训练数据有限的问题,本项目将开发新的方法来自动生成不同语言的标记数据。该项目将引入技术来减轻生成数据中的噪音,并优化生成程序,以促进多语言学习和性能。该项目的研究活动将与教育和外展任务紧密结合,以扩大其影响。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Natural language text is replete with important events in different areas (protests, cybersecurity breaches, elections, disease outbreaks, and business transactions). Identifying events to describe who did what to whom and their relations (causal, subevent, and coreferential) from a large amount of text can provide valuable data to support intelligent applications and data-driven decisions over various domains. However, current event structure extraction systems can only perform over text data for a few popular languages such as English, Chinese, Spanish, and Arabic. Text data from many other languages in the world thus cannot be processed by current event extraction systems. This limitation has hindered the coverage of data sources for the systems, introduced language biases in the extracted events, and delayed updates with latest events in local reports. Eventually, the collected event data from current techniques cannot comprehensively represent the latest dynamics over the world to effectively support decision making for important problems of national interests. To address the multilingual challenges, this project will develop event extraction and event-event relation extraction systems that can be effective for data in multiple languages, emphasizing on understudied and low-resource languages to improve the coverage of extracted data and promote democratization of technologies. In information retrieval, multilingual event structure data from the developed technologies can enable data management systems to quickly obtain answers and create summaries for broader user queries in many more languages. In cybersecurity, databases for extracted cyber attack events from multilingual sources can be used to generate more fine-grained and comprehensive reports to inform resource allocation decisions to better protect online activities. In socio-political science, coded conflict and meditation events from more languages can increase the scope and reduce biases of the data to support better decisions for foreign policy, civil war prevention, environmental challenges, or economic strategies.This project will address three fundamental limitations of existing multilingual learning research for event structure extraction: (i) the lack of multilingual datasets that provide data annotation for multiple languages to sufficiently support generalization evaluation of models across different language families, (ii) the limitations of current multilingual representation learning methods when aligning representations between languages to induce language-general features, and (iii) the scarcity of labeled data in different languages to train multilingual models. First, the project will annotate documents for all event extraction and event-event relation extraction tasks in many more languages using consistent schemas. The selected languages for annotation will be typologically diverse, understudied and low-resource to provide reliable multilingual evaluation data for the developed methods. Second, to boost cross-lingual performance for event structure extraction, this project will devise multilingual representation learning methods to enable effective knowledge transfer where models trained on labeled data of high-resource languages can be directly applied to data of other languages. The project will develop novel representation alignment methods for different languages using representation matching, augmentation, and language-general structure induction for text. Third, concerning limited training data for multilingual learning, this project will develop novel methods to automatically generate labeled data in different languages. The project will introduce techniques to mitigate noises in the generated data and optimize generation procedures to boost multilingual learning and performance. The research activities in this project will be closely integrated with education and outreach missions to broaden their impacts.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.18653/v1/2023.findings-acl.727
发表时间:
2023
期刊:
影响因子:
--
作者:
[Amir Pouran Ben Veyseh;Franck Dernoncourt;Bonan Min;Thien Huu Nguyen]
通讯作者:
Amir Pouran Ben Veyseh;Franck Dernoncourt;Bonan Min;Thien Huu Nguyen
DOI:
10.18653/v1/2023.findings-acl.135
发表时间:
2023
期刊:
影响因子:
--
作者:
[Chien Nguyen;Linh Van Ngo;Thien Huu Nguyen]
通讯作者:
Chien Nguyen;Linh Van Ngo;Thien Huu Nguyen
DOI:
10.18653/v1/2023.acl-long.296
发表时间:
2023
期刊:
影响因子:
--
作者:
[Luis Guzman Nateras;Franck Dernoncourt;Thien Huu Nguyen]
通讯作者:
Luis Guzman Nateras;Franck Dernoncourt;Thien Huu Nguyen
DOI:
10.18653/v1/2023.findings-acl.266
发表时间:
2023
期刊:
影响因子:
--
作者:
[Chien Van Nguyen;Hieu Man;Thien Huu Nguyen]
通讯作者:
Chien Van Nguyen;Hieu Man;Thien Huu Nguyen
Phase I IUCRC University of Oregon: Center for Big Learning
-
批准号:1747798
-
项目类别:Continuing Grant
-
资助金额:$75.0万
-
财政年份:2018
-
负责人:Thien Nguyen
-
依托单位:
海外基金