课题基金 / 基金详情

CAREER: Towards Open World Event Knowledge Extraction with Weak Supervision

CAREER: Towards Open World Event Knowledge Extraction with Weak Supervision
职业:在弱监督下实现开放世界事件知识提取
批准号:
2238940
负责人:
Lifu Huang
金额:
$59.35万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-08-15 至 2028-07-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
了解事件,比如谁在何时何地对谁做了什么,是人类了解不断变化的世界的基本活动之一。这些问题的答案是绝大多数(如果不是全部的话)基于语言的交流所传达的关键信息的基础。然而,目前的研究范式在从开放世界场景中提取事件知识方面存在一些不足。在这些场景中,从数据中提取知识仅限于几个大领域(例如,新闻或生物医学)或通用语言(例如,英语,西班牙语和中文),因为严重依赖于人类对数据进行上下文化的努力。这包括为一些目标事件类型创建大规模的手动注释或定义示意图模板。本项目旨在通过开发新的、更高效的算法,为开放世界事件知识提取奠定基础,建立新的范式,将提取能力扩展到更广泛的场景,同时需要最少的人力。这个基础应该提供不同事件类型的广泛覆盖,并且很容易适应新出现的场景。该项目的成功将直接惠及智能信息接入系统的用户。对于分析新兴和趋势主题和事件的应用程序,例如自然灾害,全国选举,抗议和疾病爆发,提议的研究的成功不仅将为人类提供准确和抽象的总结和易于访问的每个主题,而且还可以让分析人员更好地发现事件的参与者,原因,影响和时间顺序,并帮助发现更多的见解。该项目的技术目标分为三个重点。Thrust 1开发了模式引导的事件提取方法。这是通过利用来自复杂目标事件模式的知识来完成的,例如事件类型结构(即类型名称和参数角色)、层次结构和事件类型之间的时间/因果/部分-整体关系,这些知识提供了有价值的指导,特别是在很少或没有可用注释的情况下。虽然大多数领域和场景的事件注释都不存在,并且获取它们非常昂贵和耗时,但通常可以访问大规模的未标记域内数据。因此,Thrust 2将进一步开发一套更有效和新颖的自我训练策略,通过自我监督来利用大规模的未标记数据。在实践中,甚至没有事件类型模式可用于大多数域和场景,例如自然灾害或疾病爆发。手动定义具有高覆盖率的事件模式是极具挑战性和耗时的,因为它需要语言学和目标领域的背景知识,并且人类需要手动检查大量的领域内数据以确定突出的事件类型。考虑到这些挑战,Thrust 3进一步探索了新的解决方案,以从原始文本中自动推断目标事件模式,包括事件类型、参与者的角色以及它们之间的关系,并相应地提取它们的事件提及。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Understanding events, such as who did what to whom, when and where, is one of the fundamental human activities to learn about the changing world. The answers to these questions underpin the key information conveyed in the overwhelming majority, if not all, of language-based communication. However, current research paradigm suffers from several shortcomings in extracting event knowledge from the open world scenarios. In these scenarios, knowledge extraction from data is limited to a few large domains (e.g., news or biomedical) or common languages (e.g., English, Spanish and Chinese), because of the heavy reliance on the human effort to contextualize data. This includes creating large-scale manual annotations or defining the schematic templates for a few target event types. This project aims to lay the foundation and establish new paradigms for open world event knowledge extraction by developing new and more efficient algorithms to extend the extraction capability to the wide range of scenario, while requiring minimal human effort. This foundation should provide extensive coverage of different event types and be easily adapted to emerging scenarios. The success of this project will directly benefit users of the intelligent information access systems. For applications that analyze emerging and trending topics and events, such as natural disasters, national elections, protest and disease outbreak, success of the proposed research will not only provide an accurate and abstractive summary and easy access of each topic for humans, but also allow analysts to better discover the participants of the events, the cause, effects and temporal orders among them, and help discover more insights. The technical aims of the project are divided into three thrusts. Thrust 1 develops schema-guided event extraction approaches. This is done by leveraging the knowledge from the complex target event schema, such as the event type structures (i.e., type name and argument roles), hierarchy and temporal/causal/part-whole relations among the event types, which provide valuable guidance, especially when there is few to no annotations available. While event annotations for most of the domains and scenarios are not existing and extremely expensive and time-consuming to obtain, the large-scale unlabeled in-domain data are usually accessible. Thus, Thrust 2 will further develops a suite of more efficient and novel self-training strategies to make use of the large-scale unlabeled data through self-supervision. In practice, there is even no event type schema available to most of the domains and scenarios, such as natural disaster or disease outbreak. Manually defining an event schema with high coverage is extremely challenging and time consuming as it requires background knowledge in both linguistics and the target domain, and humans need to manually examine a large amount of in-domain data to determine the salient event types. Considering these challenges, Thrust 3 further explores novel solutions to automatically deduce the target event schema, including event types, the roles of their participants, as well as their relations from the raw text and extract their event mentions accordingly.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金