课题基金 / 基金详情

Using case difficulty to improve predictive performance evaluation

Using case difficulty to improve predictive performance evaluation
利用案例难度来改进预测绩效评估
批准号:
RGPIN-2021-02588
负责人:
Lee, Joon
金额:
$2.11万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31

项目摘要

项目成果

Lee, Joon的其他基金

相似基金

相关文献

中文摘要
翻译
近年来,机器学习(ML)在人工智能(AI)领域带来了许多突破性的创新。特别是,预测模型一直处于这场人工智能革命的最前沿(例如,照片中的物体检测)。尽管看起来取得了巨大而迅速的进展,但机器学习的某些方面几十年来并没有太大的变化。其中一个停滞不前的领域是如何在预测性绩效评估中平等对待所有案例。大多数性能指标仅仅基于预测正确或错误的情况,而完全忽略了哪些情况。因此,没有考虑案例难度的概念(即正确预测给定案例的难度有多大?),这通常会导致误导性的性能评估。尽管在许多预测问题中,不同的情况往往表现出不同程度的困难,但正确/错误预测困难情况的权重与正确/错误预测简单情况的权重相同。一个相关的问题是是否所有可用的测试用例都需要用于预测性能评估。仅仅使用基于案例难度选择的少量测试用例,是否有可能准确地评估性能?这种方法的灵感来自于计算机化的考试,比如研究生入学考试(GRE),它只选择和使用一组问题中的一个子集。这种方法可以防止绩效评估被太容易或太难的案例所主导,并减轻许多问题领域中数据不足的问题。我的长期研究计划侧重于开发横切多个问题域的新型机器学习方法。在接下来的5年里,我的研究将集中在开发新的预测性能指标和评估方法。目标是:1。根据案例难度开发新的预测性能指标。这些性能指标将确保对困难情况的正确预测得到奖励,而对简单情况的错误预测将受到惩罚。2.开发新的基于用例难度的预测性能评估方法,比传统方法需要更少的测试用例。测试用例将按顺序呈现,下一个测试用例将根据用例难度和之前的测试用例是否被正确预测来选择。这项研究的主要交付成果将是一个开源的Python包,它实现了我们新颖的预测性能指标和评估方法。该软件包将使机器学习研究人员和工程师能够将案例难度纳入他们的预测建模,并精确评估他们的模型。在我们这个日益由人工智能驱动的社会,预测模型比以往任何时候都更加普遍。这项研究很重要,因为它将使我们全面了解这些影响我们生活的预测模型是如何真正发挥作用的。加拿大人将受益于由此产生的可信赖的预测模型,进一步加强加拿大在机器学习研究方面的国际领导地位。
英文摘要
In recent years, machine learning (ML) has led to a number of ground-breaking innovations in artificial intelligence (AI). In particular, predictive models have been at the forefront of this AI revolution (e.g., object detection in photos). Despite this seemingly substantial and rapid progress, some aspects of ML have not changed much for decades. One such area of stagnation is how all cases are treated equally in predictive performance evaluation. Most performance metrics are solely based on how many cases are predicted correctly or incorrectly and completely ignore which cases. As a result, the concept of case difficulty (i.e., how difficult is it to predict a given case correctly?) is not considered, often resulting in misleading performance evaluation. Although different cases tend to exhibit varying degrees of difficulty in many prediction problems, correctly/incorrectly predicting a difficult case is weighted just the same as correctly/incorrectly predicting an easy case. A related issue is whether all available test cases need to be used for predictive performance evaluation. Is it possible to accurately assess performance by only using a small number of test cases selected based on case difficulty? Inspired by how computerized tests such as the Graduate Record Examination (GRE) select and use only a subset of a pool of questions, this approach can prevent performance evaluation from being dominated by too easy or too difficult cases and mitigate insufficient data issues in many problem domains. My long-term research program focuses on developing novel ML methodologies that crosscut multiple problem domains. Over the next 5 years, my research will focus on developing novel predictive performance metrics and evaluation methods. The objectives are to: 1.Develop new predictive performance metrics based on case difficulty. These performance metrics will ensure correct prediction of difficult cases is rewarded, whereas incorrect prediction of easy cases is penalized. 2.Develop new case difficulty-based predictive performance evaluation methods that require fewer test cases than traditional methods. Test cases will be presented sequentially and the next test case will be selected based on case difficulty and whether the previous test cases were predicted correctly. The primary deliverable of this research will be an open-source Python package that implements our novel predictive performance metrics and evaluation methods. This package will enable ML researchers and engineers to incorporate case difficulty into their predictive modeling and precisely evaluate their models. In our increasingly AI-driven society, predictive models are more prevalent than ever. This research is important because it will lead to comprehensive insights into how these predictive models that affect our lives truly perform. Canadians will benefit from resulting trustworthy predictive models, further strengthening Canada's international leadership in ML research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Using case difficulty to improve predictive performance evaluation
  • 批准号:
    RGPIN-2021-02588
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.11万
  • 财政年份:
    2022
  • 负责人:
    Lee, Joon
  • 依托单位:
Personalized Decision Support Driven by Similarity Metrics
  • 批准号:
    RGPIN-2014-04743
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2019
  • 负责人:
    Lee, Joon
  • 依托单位:
Personalized Decision Support Driven by Similarity Metrics
  • 批准号:
    RGPIN-2014-04743
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2018
  • 负责人:
    Lee, Joon
  • 依托单位:
Personalized Decision Support Driven by Similarity Metrics
  • 批准号:
    RGPIN-2014-04743
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2017
  • 负责人:
    Lee, Joon
  • 依托单位:
国内基金
海外基金
Intelligent Patent Analysis for Optimized Technology Stack Selection:Blockchain BusinessRegistry Case Demonstration
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    USHARANI HAREESH GOVINDARA JAN
  • 依托单位:
一类特殊Abelian群的子群计数问题
  • 批准号:
    12301006
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30.00万元
  • 批准年份:
    2023
  • 负责人:
    隋延坤
  • 依托单位:
联合CRISPR技术以及iPSC技术区别研究AMD高危序列ARMS2以及HTRA1基因型致病机理
  • 批准号:
    81670875
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2016
  • 负责人:
    李筱荣
  • 依托单位:
Case-Cohort数据的半参数逆回归估计和纵向数据分析
  • 批准号:
    11071137
  • 项目类别:
    面上项目
  • 资助金额:
    22.0万元
  • 批准年份:
    2010
  • 负责人:
    杨瑛
  • 依托单位: