III: Small: Revisiting Experimental Evaluation Protocols for Link Prediction in Knowledge Graphs
III: Small: Revisiting Experimental Evaluation Protocols for Link Prediction in Knowledge Graphs
批准号:
2346959
负责人:
Carlos Rivero
金额:
$39.35万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-06-01 至 2027-05-31
中文摘要
该项目旨在提高对知识图中链接预测的理解。知识图谱通过链接连接信息。例如,一个人被链接到一部电影,因为他们在电影中扮演角色。这允许将此人链接到电影中的其他演员,或将电影链接到该演员的其他电影。这些链接创建了一个知识图谱。许多服务,如搜索引擎,越来越依赖知识图谱,导致每天有数百万用户与这些图谱互动。知识图谱通常是不完整的;也就是说,在实际上相关的实体之间有许多缺失的链接。这些缺失的环节阻碍了图表的有效性。例如,缺少链接的搜索引擎不能完整或准确地回答用户问题。链接预测算法旨在使知识图更完整,从而更准确。当前的评估协议忽略了由链接预测算法预测的链接的性质(可解释性),并且依赖于使用随机选择和任意阈值(偏差)制作的数据集。因此,目前对链接预测算法的优点和缺点的了解相当有限。如果没有这样的理解,知识图中的链接预测领域就不能很好地向前发展,因为没有明确的方向去做,链接预测算法通常依赖于机器学习,所以他们训练链接预测模型来完成知识图。有三个主要问题阻碍了我们对链接预测模型的理解:1)缺乏解释一组链接预测而不是单个预测的方法,以及量化模型的可解释性;2)使用随机选择和任意阈值来评估引入偏差的链接预测;以及3)由于实验评估协议的变化而缺乏同类比较,从而阻碍了可复制性。该项目提出了以下进展:1)计算模型认为正确的链接预测的全局解释的新方法;2)检测和解释模型已经学习的链接预测规则的新方法,例如,“如果演员A在电影M中表演,那么A在M的演员阵容中”;3)新的解释度量和分析,考虑了各种策略来生成不正确的知识及其预期的似真度;4)异常的新定义,也就是。数据冗余,在考虑链接预测规则的同时了解基准数据集中的偏差;5)将数据集划分为保留图表特征并提供统计保证的新方法;6)从真实知识图中选择子图作为基准数据集的新统计可靠方法;7)开源链接预测框架,以减少复制结果时的障碍;8)可通过Google Colab获得的链接预测评估模块,以及用于促进开放比较的公开链接预测模型。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project aims to advance the understanding of link prediction in knowledge graphs. Knowledge graphs connect information through links. For example, a person is linked to a movie because they acted in the movie. This allows that person to be linked to other actors in the movie or to link the movie to the actor's other movies. These links create a knowledge graph. Many services like search engines increasingly rely on knowledge graphs, resulting in millions of users interacting with these graphs daily. Knowledge graphs are typically incomplete; that is, there are many missing links between entities that are in fact related. These missing links hinder graphs' effectiveness. For example, a search engine with missing links cannot completely or accurately answer a user question. Link prediction algorithms aim to make knowledge graphs more complete and therefore more accurate. Current evaluation protocols ignore the nature of the links predicted by a link prediction algorithm (interpretability), and rely on datasets crafted using random selection and arbitrary thresholds (bias). The current understanding of benefits and drawbacks of link prediction algorithms is thus quite limited. Without such understanding, the field of link prediction in knowledge graphs cannot properly advance as there is no clear direction on how to do so.Link prediction algorithms commonly rely on machine learning, so they train link prediction models to complete knowledge graphs. There are three main issues hindering our understanding of what link prediction models can accomplish: 1) The lack of methods to interpret a set of link predictions rather than individual predictions, and to quantify model interpretability; 2) The use of random selection and arbitrary thresholds to evaluate link prediction that introduce biases; and 3) The lack of homogeneous comparisons that hinder replicability due to variations in the experimental evaluation protocol. This project proposes the following advances: 1) New methods to compute global interpretations of the link predictions that a model deems correct; 2) New methods to detect and interpret the link prediction rules a model has learned, such as "if actor A acts in movie M, then A is in M's cast;" 3) New interpretation metrics and analyses considering various strategies to generate incorrect knowledge and its expected plausibility; 4) New definitions of anomalies, a.k.a. data redundancy, to understand biases in benchmarking datasets while taking link prediction rules into account; 5) New methods to partition datasets into splits that preserve graph features with statistical guarantees; 6) New statistically sound methods to select subgraphs from real-world knowledge graphs for use as benchmarking datasets; 7) An open-source link prediction framework to reduce barriers when replicating results; and 8) A link prediction evaluation module available through Google Colab, and publicly-available link prediction models to promote open comparisons.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Using Program Dependence Graphs to Propagate Feedback to Students on Programming Assignments and Promote Responsive Teaching
-
批准号:1915404
-
项目类别:Standard Grant
-
资助金额:$29.87万
-
财政年份:2019
-
负责人:Carlos Rivero
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: