III: Medium: Collaborative Research: Mining and Leveraging Knowledge Hypercubes for Complex Applications
III: Medium: Collaborative Research: Mining and Leveraging Knowledge Hypercubes for Complex Applications
批准号:
1956151
负责人:
Jiawei Han
金额:
$40.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2024-09-30
中文摘要
知识库是指一种机器可读的结构,它存储了关于各种实体(如组织、事件、基因)的知识,便于高效地查找信息。在许多领域中,知识随上下文而变化,现有知识库通常采用的平面结构无法捕获与不同上下文相关的复杂知识。为了使知识资源更易于查找、访问、互操作和重用(FAIR),该项目计划概念化一种新的结构,即知识超立方体,用于组织和检索可以支持各个领域复杂应用程序的知识。知识超立方体根据选定的重要维度(如时间、地点、条件)对知识进行组织,从而使人们能够在任何情况下轻松获取知识,封装独特的实体和事实,并进行跨维度的比较和推理。该项目影响了人们发现和使用知识的方式,推进了基于知识的数据分析方法,并通过在它们之间建立桥梁,使具有大量文献和未解决的复杂任务的广泛领域受益。知识超多维数据集还可以支持教育创新,并有助于知识追踪等教育任务。本提案的主要目标是形成一个从大量文本文档中挖掘知识混合数据集的范例,并利用这些混合数据集进行复杂的探索和预测任务。为了实现这一目标,该项目解决了一系列技术挑战。首先,为了从海量文本中自动构建知识超立方体,设计了创新的弱监督方法,基于超立方体结构组织文本文档,提取开放实体和关系信息,多维地组织特定于细胞和跨细胞的知识。其次,开发了新的改进方法,通过在超立方体内部和外部信息进行交叉检查,自动验证知识超立方体中单元内和单元间的信息质量。第三,知识超立方体推动了新的发现和学习任务的发展。特别是,该项目引入了一个自动知识搜索管道,用于利用知识超立方体进行下游预测任务,以及一个假设生成方法,用于对概念之间的未知关联进行评分。规划的范式在两个特定领域(即生物医学和新闻事件)中实现,展示了知识超立方体的力量,可以对这些领域产生新的见解。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Knowledge repository refers to a machine-readable structure that stores knowledge about various entities (e.g., organizations, events, genes), which facilitates efficient information seeking. In many domains, knowledge varies with respect to contexts, and a flat structure that is commonly adopted by existing knowledge repositories cannot capture the complicated knowledge associated with different contexts. To make knowledge resources more findable, accessible, interoperable, and reusable (FAIR), this project plans to conceptualize a new structure, Knowledge Hypercube, for organizing and retrieving knowledge that could support complex applications in various domains. A knowledge hybercube organizes knowledge with respect to selected important dimensions (e.g., time, locations, conditions), and thus it allows people to easily access knowledge in any context, encapsulate distinctive entities and facts, and conduct cross-dimensional comparison and inference. This project impacts how people find and use knowledge, advances knowledge-based data analytics approaches, and benefits a wide range of domains which have gigantic literature and unsolved complex tasks by building a bridge between them. Knowledge hypercubes can also support educational innovation and contributes to educational tasks such as knowledge tracing. The major objective of this proposal is to form a paradigm of mining knowledge hybercubes from massive collection of text documents and leveraging such hybercubes for complex exploration and prediction tasks. To meet this goal, this project tackles a series of technical challenges. First, to automatically construct a knowledge hypercube from massive texts, innovative weakly supervised approaches are designed to organize text documents based on the hypercube structure, extract open entity and relationship information and organize cell-specific and cross-cell knowledge in a multi-dimensional manner. Second, novel refinement approaches are developed to automatically verify the information quality within and across cells in knowledge hypercubes by cross-checking within the hypercubes and with external information. Third, knowledge hypercubes motivate the development towards new discovery and learning tasks. In particular, the project introduces an automatic knowledge search pipeline for leveraging knowledge hypercubes for downstream prediction tasks, and a hypothesis generation approach for scoring unknown associations between concepts. The planned paradigm is realized in two specific domains (i.e., biomedical and news events), demonstrating the power of knowledge hypercubes to enable new insights into these domains.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(40)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1109/bigdata50022.2020.9378031
发表时间:
2020-12
期刊:
2020 IEEE International Conference on Big Data (Big Data)
影响因子:
--
作者:
[Carl Yang;Liyuan Liu;Mengxiong Liu;Zongyi Wang;Chao Zhang;Jiawei Han]
通讯作者:
Carl Yang;Liyuan Liu;Mengxiong Liu;Zongyi Wang;Chao Zhang;Jiawei Han
Corpus-Based Relation Extraction by Identifying and Refining Relation Patterns
通过识别和细化关系模式进行基于语料库的关系提取
DOI:
--
发表时间:
2023
期刊:
Springer
影响因子:
--
作者:
[Sizhe Zhou, Suyu Ge]
通讯作者:
Sizhe Zhou, Suyu Ge
DOI:
10.1145/3442381.3450114
发表时间:
2021-02
期刊:
Proceedings of the Web Conference 2021
影响因子:
--
作者:
[Xinyang Zhang;Chenwei Zhang;Xin Dong;Jingbo Shang;Jiawei Han]
通讯作者:
Xinyang Zhang;Chenwei Zhang;Xin Dong;Jingbo Shang;Jiawei Han
DOI:
10.18653/v1/2023.findings-acl.14
发表时间:
2023
期刊:
Blood
影响因子:
20.3
作者:
[Nishant Balepur;Shivam Agarwal;Karthik Venkat Ramanan;Susik Yoon;Diyi Yang;Jiawei Han]
通讯作者:
Nishant Balepur;Shivam Agarwal;Karthik Venkat Ramanan;Susik Yoon;Diyi Yang;Jiawei Han
Unsupervised Story Discovery from Continuous News Streams via Scalable Thematic Embedding
通过可扩展的主题嵌入从连续新闻流中无监督地发现故事
DOI:
10.1145/3539618.3591782
发表时间:
2023
期刊:
ACM
影响因子:
--
作者:
[Yoon, Susik, Lee, Dongha, Zhang, Yunyi, Han, Jiawei]
通讯作者:
Han, Jiawei
共 36 条
BIGDATA: F: Collaborative Research: Taming Big Networks via Embedding
-
批准号:1741317
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2018
-
负责人:Jiawei Han
-
依托单位:
III: Medium: Collaborative Research: StructNet: Constructing and Mining Structure-Rich Information Networks for Scientific Research
-
批准号:1704532
-
项目类别:Continuing Grant
-
资助金额:$40.0万
-
财政年份:2017
-
负责人:Jiawei Han
-
依托单位:
III: Small: Multi-Dimensional Structuring, Summarizing and Mining of Social Media Data
-
批准号:1618481
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2016
-
负责人:Jiawei Han
-
依托单位:
III: Small: Collaborative Research: Conflicts to Harmony: Integrating Massive Data by Trustworthiness Estimation and Truth Discovery
-
批准号:1320617
-
项目类别:Continuing Grant
-
资助金额:$21.11万
-
财政年份:2013
-
负责人:Jiawei Han
-
依托单位:
III-Core:Small: MoveMine: Mining Sophisticated Patterns and Actionable Knowledge from Massive Moving Object Data
-
批准号:1017362
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2010
-
负责人:Jiawei Han
-
依托单位:
CPS: Small: Collaborative Research: Foundations of Cyber-Physical Networks
-
批准号:0931975
-
项目类别:Standard Grant
-
资助金额:$27.5万
-
财政年份:2009
-
负责人:Jiawei Han
-
依托单位:
III: Medium: Collaborative Research: Towards On-Line Analytical Mining of Heterogeneous Information Networks
-
批准号:0905215
-
项目类别:Standard Grant
-
资助金额:$83.13万
-
财政年份:2009
-
负责人:Jiawei Han
-
依托单位:
SGER: CS-BibCube: OLAPing and Mining of Computer Science Literature
-
批准号:0842769
-
项目类别:Standard Grant
-
资助金额:$10.5万
-
财政年份:2008
-
负责人:Jiawei Han
-
依托单位:
SGER: DataScope: Viewing Database Contents in Multi-Resolution at Your Finger Tips
-
批准号:0642771
-
项目类别:Standard Grant
-
资助金额:$8.0万
-
财政年份:2006
-
负责人:Jiawei Han
-
依托单位:
Collaborative Research: Endowing Biological Databases With Analytical Power: Indexing, Querying, and Mining of Complex Biological Structures
-
批准号:0515813
-
项目类别:Standard Grant
-
资助金额:$27.14万
-
财政年份:2005
-
负责人:Jiawei Han
-
依托单位:
SEI(IIS): MotionEye: Querying and Mining Large Datasets of Moving Objects
-
批准号:0513678
-
项目类别:Standard Grant
-
资助金额:$27.5万
-
财政年份:2005
-
负责人:Jiawei Han
-
依托单位:
NGS: Collaborative Research: Reusable, Observation-based Performance Prediction across Platforms
-
批准号:0406408
-
项目类别:Standard Grant
-
资助金额:$2.34万
-
财政年份:2004
-
负责人:Jiawei Han
-
依托单位:
Mining Dynamics of Data Streams in Multi-Dimensional Space
-
批准号:0308215
-
项目类别:Continuing Grant
-
资助金额:$30.0万
-
财政年份:2003
-
负责人:Jiawei Han
-
依托单位:
Mining Sequential Patterns and Structured Patterns: Scalability, Flexibility, Extensibility, and Applicability
-
批准号:0209199
-
项目类别:Continuing Grant
-
资助金额:$16.5万
-
财政年份:2002
-
负责人:Jiawei Han
-
依托单位:
海外基金