III: Small: Revisiting Experimental Evaluation Protocols for Link Prediction in Knowledge Graphs
III: Small: Revisiting Experimental Evaluation Protocols for Link Prediction in Knowledge Graphs
批准号:
2346959
负责人:
Carlos Rivero
金额:
$39.35万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-06-01 至 2027-05-31
中文摘要
该项目旨在促进对知识图中链接预测的理解。知识图谱通过链接连接信息。例如,一个人被链接到一部电影,因为他们在电影中扮演角色。这允许该人被链接到电影中的其他演员或将电影链接到演员的其他电影。这些链接创建了一个知识图谱。搜索引擎等许多服务越来越依赖于知识图谱,导致每天有数百万用户与这些图谱进行交互。知识图谱通常是不完整的;也就是说,在实际上相关的实体之间存在许多缺失的链接。这些缺失的环节阻碍了图表的有效性。例如,缺少链接的搜索引擎无法完整或准确地回答用户的问题。链接预测算法旨在使知识图更完整,因此更准确。目前的评估协议忽略了链接预测算法预测的链接的性质(可解释性),并依赖于使用随机选择和任意阈值(偏差)制作的数据集。因此,目前对链路预测算法的优点和缺点的理解是相当有限的。如果没有这样的理解,知识图中的链接预测领域就无法正常发展,因为没有明确的方向如何做到这一点。链接预测算法通常依赖于机器学习,因此它们训练链接预测模型来完成知识图。有三个主要问题阻碍了我们对链接预测模型的理解:1)缺乏解释一组链接预测而不是单个预测的方法,以及量化模型可解释性的方法; 2)使用随机选择和任意阈值来评估引入偏差的链接预测;和3)由于实验评估方案的变化,缺乏阻碍可重复性的同质比较。该项目提出了以下进展:1)计算模型认为正确的链接预测的全局解释的新方法; 2)检测和解释模型学习的链接预测规则的新方法,例如“如果演员A在电影M中表演,那么A在M的演员阵容中;“3)新的解释指标和分析,考虑到各种策略,以产生不正确的知识及其预期的可解释性;数据冗余,以了解基准数据集的偏差,同时考虑到链接预测规则; 5)将数据集划分为具有统计保证的图特征的分割的新方法; 6)从真实世界的知识图中选择子图用作基准数据集的新统计合理方法; 7)开源链接预测框架,以减少复制结果时的障碍;和8)通过Google Colab提供的链接预测评估模块,以及公开提供的链接预测模型,以促进开放式比较。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project aims to advance the understanding of link prediction in knowledge graphs. Knowledge graphs connect information through links. For example, a person is linked to a movie because they acted in the movie. This allows that person to be linked to other actors in the movie or to link the movie to the actor's other movies. These links create a knowledge graph. Many services like search engines increasingly rely on knowledge graphs, resulting in millions of users interacting with these graphs daily. Knowledge graphs are typically incomplete; that is, there are many missing links between entities that are in fact related. These missing links hinder graphs' effectiveness. For example, a search engine with missing links cannot completely or accurately answer a user question. Link prediction algorithms aim to make knowledge graphs more complete and therefore more accurate. Current evaluation protocols ignore the nature of the links predicted by a link prediction algorithm (interpretability), and rely on datasets crafted using random selection and arbitrary thresholds (bias). The current understanding of benefits and drawbacks of link prediction algorithms is thus quite limited. Without such understanding, the field of link prediction in knowledge graphs cannot properly advance as there is no clear direction on how to do so.Link prediction algorithms commonly rely on machine learning, so they train link prediction models to complete knowledge graphs. There are three main issues hindering our understanding of what link prediction models can accomplish: 1) The lack of methods to interpret a set of link predictions rather than individual predictions, and to quantify model interpretability; 2) The use of random selection and arbitrary thresholds to evaluate link prediction that introduce biases; and 3) The lack of homogeneous comparisons that hinder replicability due to variations in the experimental evaluation protocol. This project proposes the following advances: 1) New methods to compute global interpretations of the link predictions that a model deems correct; 2) New methods to detect and interpret the link prediction rules a model has learned, such as "if actor A acts in movie M, then A is in M's cast;" 3) New interpretation metrics and analyses considering various strategies to generate incorrect knowledge and its expected plausibility; 4) New definitions of anomalies, a.k.a. data redundancy, to understand biases in benchmarking datasets while taking link prediction rules into account; 5) New methods to partition datasets into splits that preserve graph features with statistical guarantees; 6) New statistically sound methods to select subgraphs from real-world knowledge graphs for use as benchmarking datasets; 7) An open-source link prediction framework to reduce barriers when replicating results; and 8) A link prediction evaluation module available through Google Colab, and publicly-available link prediction models to promote open comparisons.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Using Program Dependence Graphs to Propagate Feedback to Students on Programming Assignments and Promote Responsive Teaching
-
批准号:1915404
-
项目类别:Standard Grant
-
资助金额:$29.87万
-
财政年份:2019
-
负责人:Carlos Rivero
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: