课题基金 / 基金详情

Exploring Trust in AI Enabled Systems

Exploring Trust in AI Enabled Systems
探索人工智能系统的信任
批准号:
2407974
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
虽然评估人工智能系统中可解释技术的有效性的方法多种多样,但可解释性或可解释性的度量(缺乏通用的分类)很少考虑信任问题。Pieters(2010)和Miller(2019)之前的研究表明,这种关系并不像看起来那么微不足道,因为,例如,完整但不健全的解释会改善用户的心理模型,但会降低信任,而细节水平过高的解释根本无法引起任何信任。本研究的目的是探索衡量人工智能系统信任的更有效方法,并将研究范围扩展到服务关系中认知信任和情感信任的区别。前者是用户从积累的知识中对服务提供者的能力和可靠性产生的信心或意愿,而后者则是用户对实体的信心,这种信心是基于实体所表现出的关心和关怀水平所产生的感觉。基于这一区别,是否应该对人工智能系统(如自主代理)的安全性和安全性中的用户信任进行评估,采用情感信任文献中的技术?Rempel et al.(1985)特别提到了自动化中的信任,他声称信任与可预测性、可靠性和信念这三个维度一起发展。然而,其他著作似乎不同意这些维度是什么,指出试错经验、理解和信仰(Zuboff, 1988),可靠性和信仰(McKnight等人,2002),以及经验、理解和可观察性(Rogers, 2003)。因此,需要考虑多个因素,包括机器学习系统可能对谁可解释的问题。在现有的可解释性文献中缺少的一个与安全和安保特别相关的利益相关者是攻击者。在此,进一步的研究将集中于在模型内部建立不确定性意识,以提高系统的鲁棒性。然而,由于基于透明度的解释技术比事后技术更容易被攻击者利用,因此本研究还将在权衡可解释性的背景下探索和评估XAI技术本身的安全性。该项目采取了许多方向,计划评估和比较当前的XAI技术,并设计新颖的正式技术,解决众多利益相关者对人类- ai系统信任的测量。
英文摘要
While methods of evaluating the efficacy of explainable techniques in AI systems are numerous and varied, measures of explainability or interpretability (a common taxonomy is lacking) rarely consider the issue of trust. Previous studies by Pieters (2010) and Miller (2019) show that this relationship is not as trivial as it seems, as, for example, complete but unsound explanations improve the user's mental model but reduce trust, while explanations with a level of detail too high fail to elicit any trust at all. The aim of the study is to explore more effective ways of measuring trust in AI systems and extending the scope of the research to the distinction between cognitive and affective trust in service relationships. While the first is a user's confidence or willingness to rely on a service provider's competence and reliability derived from accumulated knowledge, the second can be described as the confidence one places in an entity based on feelings generated by the level of care and concern the entity demonstrates. Based on this distinction, should the evaluation of user trust in the safety and security of an artificially intelligent system such as an autonomous agent adopt techniques from the literature on affective trust? Specific mentions to trust in automation is found in Rempel et al. (1985), who claim that trust evolves alongside the three dimensions of predictability, dependability, and faith. However, other works seem to disagree on what these dimensions are, pointing to trial-and-error experience, understanding, and faith (Zuboff, 1988), dependability, and faith (McKnight et al., 2002), and experience, understandability, and observability (Rogers, 2003). Therefore there are multiple factors to take into account, including the question of to whom a machine learning system might be interpretable. One stakeholder particularly relevant to safety and security, which is missing in the existing explainability literature, is that of attackers. Here, further research will focus on building uncertainty awareness inside models in order to improve the robustness of the system. However, since transparency-based explanation techniques are potentially more exploitable by attackers than post-hoc techniques, the study will also explore and evaluate the security of XAI techniques themselves in the context of a trade-off with explainability. There are many directions that this project take and it is planned to evaluate and compare current XAI techniques and to design novel formal techniques addressing numerous stakeholders to the measurement of trust in human-AI systems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金