Exploring Trust in AI Enabled Systems
Exploring Trust in AI Enabled Systems
批准号:
2407974
负责人:
金额:
$0.0万
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
虽然在人工智能系统中评估可解释技术的有效性的方法多种多样,但可解释性或可解释性(缺乏通用分类)的衡量标准很少考虑信任问题。皮特斯(2010)和米勒(2019)之前的研究表明,这种关系并不像看起来那么微不足道,例如,完整但不合理的解释改善了用户的心理模型,但降低了信任,而细节水平太高的解释根本无法赢得任何信任。这项研究的目的是探索更有效的方法来衡量人工智能系统中的信任,并将研究范围扩展到服务关系中认知信任和情感信任的区分。虽然第一个是用户对依赖服务提供者的能力和可靠性的信心或意愿,这些能力和可靠性来自于积累的知识,第二个可以被描述为一个实体基于该实体所表现出的关怀和关注水平所产生的感觉而对该实体所给予的信心。基于这种区别,在评估用户对人工智能系统(如自主代理)的安全性和安全性的信任时,是否应该采用文献中关于情感信任的技术?在Rempel等人中可以找到对自动化的信任的具体提到。(1985),他声称信任伴随着可预测性、可靠性和信念这三个维度而发展。然而,其他作品似乎不同意这些维度是什么,指向试错经验、理解和信仰(Zuboff,1988)、可靠性和信仰(McKnight等人,2002),以及经验、可理解性和可观察性(Rogers,2003)。因此,有多个因素需要考虑,包括机器学习系统可能对谁解释的问题。与安全和安保特别相关的一个利益相关者是攻击者,这在现有的可解释性文献中是缺失的。在这里,进一步的研究将集中在建立模型内部的不确定性意识,以提高系统的稳健性。然而,由于基于透明度的解释技术可能比后自组织技术更容易被攻击者利用,该研究还将在与可解释性的权衡背景下探索和评估XAI技术本身的安全性。这个项目有很多方向,计划评估和比较当前的XAI技术,并设计新的正式技术,解决众多利益相关者对人类-AI系统的信任度量。
英文摘要
While methods of evaluating the efficacy of explainable techniques in AI systems are numerous and varied, measures of explainability or interpretability (a common taxonomy is lacking) rarely consider the issue of trust. Previous studies by Pieters (2010) and Miller (2019) show that this relationship is not as trivial as it seems, as, for example, complete but unsound explanations improve the user's mental model but reduce trust, while explanations with a level of detail too high fail to elicit any trust at all. The aim of the study is to explore more effective ways of measuring trust in AI systems and extending the scope of the research to the distinction between cognitive and affective trust in service relationships. While the first is a user's confidence or willingness to rely on a service provider's competence and reliability derived from accumulated knowledge, the second can be described as the confidence one places in an entity based on feelings generated by the level of care and concern the entity demonstrates. Based on this distinction, should the evaluation of user trust in the safety and security of an artificially intelligent system such as an autonomous agent adopt techniques from the literature on affective trust? Specific mentions to trust in automation is found in Rempel et al. (1985), who claim that trust evolves alongside the three dimensions of predictability, dependability, and faith. However, other works seem to disagree on what these dimensions are, pointing to trial-and-error experience, understanding, and faith (Zuboff, 1988), dependability, and faith (McKnight et al., 2002), and experience, understandability, and observability (Rogers, 2003). Therefore there are multiple factors to take into account, including the question of to whom a machine learning system might be interpretable. One stakeholder particularly relevant to safety and security, which is missing in the existing explainability literature, is that of attackers. Here, further research will focus on building uncertainty awareness inside models in order to improve the robustness of the system. However, since transparency-based explanation techniques are potentially more exploitable by attackers than post-hoc techniques, the study will also explore and evaluate the security of XAI techniques themselves in the context of a trade-off with explainability. There are many directions that this project take and it is planned to evaluate and compare current XAI techniques and to design novel formal techniques addressing numerous stakeholders to the measurement of trust in human-AI systems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金