CRII: III: Learning to Extract Events from Knowledge Base Revisions
CRII: III: Learning to Extract Events from Knowledge Base Revisions
批准号:
1464128
负责人:
Alan Ritter
金额:
$15.13万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2018-08-31
中文摘要
百科知识库(KB),如维基百科和Freebase,构成了谷歌知识图谱、Facebook图谱搜索、IBM沃森等背后的潜在智能。 这些覆盖面广的数据库包含有关实体的事实,例如一个人的雇主或一个城市的市长。 知识库不应该被简单地看作是静态的快照,然而,因为我们生活在一个不断变化的世界。 例如,选举事件可以改变一个国家的领导人,或者离婚/婚礼可以改变一个人的配偶。 今天的知识库依赖于人类编辑器来保持最新;这适用于知名实体,如名人或政治家,但手动编辑无法跟踪这些大规模知识库所涵盖的大量概念。 因此,该项目将研究持续跟踪实时文本流(包括新闻和社交媒体)的方法,并在新信息可用时自动更新知识库中的概念。 这将使新型智能系统能够不断阅读每天公开撰写的所有文本,并维护一个详细的最新知识库,描述世界的当前状态。弱监督信息提取技术的预期结果预计将有广泛的应用,包括检测Twitter上讨论的网络安全事件。该项目将为俄亥俄州州立大学及其他大学的学生提供研究培训和教育经验,因为研究成果将用于开发一个开源工具包,用于广泛分发的弱监督信息提取。当重要事件发生时,知识库贡献者经常近乎实时地编辑受影响实体的属性,例如在维基百科上。 与此同时,许多人在社交媒体和新闻中讨论这些事件。 由于改变KB实体属性的事件集很大,并且事先没有固定,因此该项目将研究,实现和评估用于从KB修订中学习文本提取器的新模型。 该项目将进行实验,学习新闻和Twitter的提取器,使用维基百科信息框编辑作为远程监督。 而不是封闭世界的假设,这是常见的在以前的工作中,所提出的方法将规范的标签分布的事件,不匹配的知识修订对用户提供的期望。 预计这项研究的结果将有助于解决由于修订历史中未反映的事件而导致的误报问题。 该方法自动提出实时维基百科信息框编辑的能力将在公众对事件的了解变得可用时进行测试。 以往的弱监督事件抽取研究大多局限于有限的领域。 相比之下,这项工作的目的是扩大规模,同时接地文本中提到的事件在知识库中的实体的属性的修订。项目网站(http://aritter.github.io/crii/)将包括项目信息、出版物链接、软件和研究成果数据集。
英文摘要
Encyclopedic knowledge bases (KBs) such as Wikipedia and Freebase form the underlying intelligence behind Google's Knowledge Graph, Facebook's Graph Search, IBM's Watson and more. These broad-coverage databases contain facts about entities, for example a person's employer or a city's mayor. KBs should not simply be viewed as static snapshots, however, as we live in a constantly changing world. For example, an election event can change the Leader of a country, or a divorce/wedding can change the Spouse of a person. Today's knowledge bases rely on human editors to stay up-to-date; this works for prominent entities, such as celebrities or politicians, but manual editing will not scale to tracking the huge number of concepts covered by these massive KBs. The project will therefore investigate methods to continuously track real-time text streams, including news and social media, and automatically update concepts in a KB, as soon as new information becomes available. This will enable new kinds of intelligent systems that constantly read all the text that is publicly written each day, and maintain a detailed up-to-the-minute knowledge base describing the current state of the world. The expected results in weakly supervised information extraction techniques are expected to have a broad range of applications, including detecting cyber security events discussed on Twitter. The project will provide research training and educational experience for students at Ohio State University and beyond, as the research outcomes will be used in developing an open-source toolkit for weakly supervised information extraction that will be widely distributed. When important events occur, KB contributors often edit properties of affected entities in near-real-time, for instance on Wikipedia. At the same time, many people discuss these events on social media and in the news. Because the set of events that alter properties of KB entities is large and not fixed in advance, this project will investigate, implement and evaluate new models for learning text extractors from KB revisions. The project will conduct experiments learning extractors for news and Twitter using Wikipedia infobox edits as distant supervision. Rather than making the closed world assumption, which is common in previous work, the proposed methods will regularize the label distribution over events that do not match knowledge revisions towards a user-provided expectation. It is expected that the results of this research will help to address the problem of false positives due to events that are not reflected in the revision history. The approach's ability to automatically propose Wikipedia infobox edits in real-time will be tested as public knowledge of an event becomes available. Previous studies on weakly supervised event extraction have mostly been conducted in limited domains. In contrast, this work aims to scale up while simultaneously grounding events mentioned in text to revisions of an entity's properties in a knowledge base. The project web site (http://aritter.github.io/crii/) will include information on the project, links to publications, software and datasets produced as a result of this research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Large-Scale Learning for Information Extraction
-
批准号:2052498
-
项目类别:Continuing Grant
-
资助金额:$48.9万
-
财政年份:2020
-
负责人:Alan Ritter
-
依托单位:
CAREER: Large-Scale Learning for Information Extraction
-
批准号:1845670
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2019
-
负责人:Alan Ritter
-
依托单位:
国内基金
海外基金
登录
查看更多内容
基于人工智能与多组学的III期结核性脓胸CT“低密度线”形成机制及手术时机预测模型研究
-
批准号:JCZRMS202602483
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:
-
依托单位:
基于MOF–CRISPR微流控平台的雄黄As(III)/As(V)价态识别与炮制耦合机制研究
-
批准号:JCZRLH202600780
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:
-
依托单位:
白术内酯III靶向IRF4-CD36轴通过调控脂质代谢重编程提升结直肠癌奥沙利铂敏感性的机制研究
-
批准号:2026JJ82690
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:张卓
-
依托单位:
基于废水零排放的FeS-As(III)置换法从污酸中清洁脱砷处理技术研究
-
批准号:2026JJ30130
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:张二军
-
依托单位:
全钒液流电池负极V(II)/V(III)电化学氧化还原的催化机理研究
-
批准号:2025JJ50094
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:王珏
-
依托单位:
猪纤维蛋白粘合剂预防胸外科术后漏气的适应症拓展研究:一项多中心、随机对照III期临床试验
-
批准号:25SF1901800
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:赵德平
-
依托单位:
HOXC8/OPN/CD44/EGFR轴介导的奥沙利铂耐药性在III期右半结肠癌耐药进展中的研究
-
批准号:2025JJ50694
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:喻南慧
-
依托单位:
MXene/nZVI@FH材料微域层界面调控水中砷(III)氧化迁移机制
-
批准号:2025JJ50319
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:陈润华
-
依托单位:
硅基III-V族亚微米线激光器的光场模式调控与耦合机理研究
-
批准号:JCZRQN202501004
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:
-
依托单位:
吡咯烷生物碱所致肝窦阻塞综合征III区肝损伤的新机制——局部氨代谢紊乱
-
批准号:JCZRYB202500652
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:
-
依托单位: