课题基金 / 基金详情

Determining Real World AI Trustworthiness and Robustness

Determining Real World AI Trustworthiness and Robustness
确定现实世界人工智能的可信度和稳健性
批准号:
10065751
负责人:
金额:
$5.1万
依托单位:
依托单位国家:
英国
项目类别:
Collaborative R&D
财政年份:
2023
资助国家:
英国
项目状态:
已结题
起止时间:
2023 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
机器学习的最新进展极大地扩展了可以自动化的任务。为改善公共服务和促进经济增长提供了巨大的机会。其中进步最大的一个领域是计算机视觉。然而,部署计算机视觉系统——特别是在安全关键环境中——是有问题的。具有机器学习的系统往往会以意想不到的方式失败,从而导致部署时的性能不达标,同时也会对使用该技术的组织造成声誉或监管损害。机器学习系统的这些失败可能是由于工程或社会期望。工程故障是指数据收集和模型训练方面的问题,如数据不平衡、模型漂移、差的分布检测、对抗性攻击等。当一个系统表现出不道德的行为时,它可能无法满足社会期望;例如,与种族或性别等受保护特征相关的糟糕预测或决策。这些故障模式的倾向应该被测量和减轻,以使AI值得信赖。在该项目中,Advai将为计算机视觉系统和指标创建沙盒测试环境,以可靠地预测这些故障模式。沙盒环境将独立于模型开发过程,并充当实际部署的代理,而无需承担部署后真正失败的风险。创建第三方测试环境的有效管道将为加速可信赖AI提供两个好处:在开发过程中提供实际性能的指示。这将大大提高效率,因为在部署时可以确定哪些模型将成功或失败;加速数据生产,随着时间的推移,确定哪些工具是部署中实际模型故障的指示。这些指标本身可以用作实际模型性能的指示器。
英文摘要
Recent advances in Machine Learning have dramatically expanded tasks that could be automated. offering enormous opportunities for improving public services and boosting economic growth. One area where advances have been largest is Computer Vision.However, deploying Computer Vision systems - particularly in safety critical environments -- is problematic. Systems with machine learning tend to fail in unexpected ways that lead to substandard performance when deployed, but also reputational or regulatory damage to the organization making use of the technology.These failures of machine learning systems can either be due to engineering or societal expectations. Engineering failures are problems with data collection and model training, such as imbalanced data, model drift, poor out-of-distribution detection, adversarial attacks etc. A system may fail to meet societal expectations when it exhibits behaviours that are unethical; for example, poor predictions or decisions related to protected characteristics such as race or gender.The propensity of these failure modes should be measured and mitigated for an AI to be trustworthy. In this project, Advai will create sandbox test environments for Computer Vision systems and metrics for reliably predicting these failure modes. The sandbox environments will be independent of the model development process and act as proxy for real-world deployment without assuming the risk of genuine failure post-deployment.An efficient pipeline for creating third party testing environments would provide two benefits for Accelerating Trustworthy AI:1.Providing an indication of real-world performance during development. This would greatly increase efficiency, as it becomes possible to identify which models will succeed or fail when deployed;2.Accelerating data production that will, over time, establish which tools are indicative of real-world model failures in deployment. These metrics themselves can be used as indicators of real-world model performance.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Immuno-Real Time PCR法精确定量血清MG7抗原及在早期胃癌预警中的价值
无色ReAl3(BO3)4(Re=Y,Lu)系列晶体紫外倍频性能与器件研究