课题基金 / 基金详情

Discovering Individual and Social Preferences through Inverse Reinforcement Learning

Discovering Individual and Social Preferences through Inverse Reinforcement Learning
通过逆强化学习发现个人和社会偏好
批准号:
ES/S00176X/1
负责人:
Amir Jahangiri
金额:
$37.1万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2018
资助国家:
英国
项目状态:
已结题
起止时间:
2018 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
提供服务和创造产品的组织通常根据调查问卷和/或与其用户群(例如患者、客户、公民)的其他明确交流形式做出决定。提供商和用户之间的这种信息交换的目的是揭示用户的“奖励功能”,即用户实际上想从他们的互动中获得什么,以及当前的产品/服务阵容存在哪些问题。为组织设计的显式信息交换形式可能既繁琐又昂贵,而且会干扰用户。此外,对于基于调查的方法来说,回答偏差是一个众所周知的问题,特别是围绕敏感话题,受访者可能出于社会或文化考虑而不愿参与。通过间接提问方法(项目计数和随机回答技术)提供了一些解决反应偏差的实用方法。然而,这些解决方案都不适用于大规模和实时设置。我们假设,理想情况下,一个组织应该尝试通过使用从用户活动中产生的观察数据来得出其用户基础的奖励函数(即用户更喜欢什么状态)。受最近人工智能研究文献的启发,我们提出了一个三方面的计划,旨在通过a)试图通过收集行为数据(例如网站点击、流量行为、电影偏好)来推断用户的奖励功能;b)创建简短、非侵入性的在线调查问卷,以消除任何不确定性;通过这项研究,我们有四个主要目标:(A)了解用户偏好,并开发方法,以通过数据和行为揭示和学习奖励功能;(B)开发交互和对话方法,以获得用户的反应和互动,从而允许用户与自动系统进行更自然的用户体验;(C)探索我们方法的社会局限性(例如,个人回报在多大程度上不是由个人偏好决定的,而是由社会强制决定的?);以及(D)调查可以采取哪些步骤,通过(A)和(B)项开发的方法获得偏好,从而完全自动化提供新服务和产品的过程。这一奖学金提供了一个独特的机会,将人工智能技术和社会科学结合在一起,以解决各种企业和组织在与客户和客户打交道时面临的问题,并试图通过行为和互动引发偏好和需求。我们将在这个项目中与我们的行业合作伙伴英国电信(BT)和埃塞克斯县议会(ECC)密切合作,调查在了解和了解他们在各自背景下面临的偏好以告知和制定工作方案方面的问题和挑战。
英文摘要
Organisations that provide services and create products often base their decisions on questionnaires and/or other explicit forms of communication with their user base (e.g. patients, customers, citizens). The aim of this information exchange between providers and users is to uncover the users' "reward function", i.e. what users actually want from their interactions and what issues exist with the current product/service line-up. Explicit forms of information exchange can be cumbersome and expensive to design for organisations and are intrusive to the user. Furthermore, response bias is a well-known problem for survey based methods, particularly around sensitive topics, where respondents maybe unwilling to engage due to social or cultural concerns. Some practical solutions to response bias are provided by indirect questioning methods (item count and randomized response techniques). However, none of these solutions are practical for large scale and real time settings.We postulate that ideally an organisation should try to elicit the reward function of its user base (i.e. what states are preferred by users) by using observational data generated from user activity. Inspired by recent literature in AI research, we propose a three-facet programme that aims to directly attack the problem of what users want by a) trying to infer the user reward function through the collection of behavioural data (e.g. website clicks, traffic behaviour, movie preferences); b) creating short, non-intrusive online questionnaires that will remove any uncertainties; and c) exploiting user preferences in order to improve service and product provision.The proposed research aims to contribute to developing methods that can be embedded in artificial intelligence systems which must elicit and understand preferences by interacting with humans in order to adapt their behaviour and allow for a more natural experience and interaction.Through this research we have four key objectives: (a) understand user preferences and develop methods to uncover and learn the reward function through data and behaviours; (b) develop interactive and conversational methods for eliciting responses and interactions from users that allow for a more natural user experience with automatic systems; (c) explore the social limitations of our approach (for instance, to what extend are personal rewards not dictated by individual preferences, but rather by social coercion?); and (d) investigate what steps can be taken to fully automate the procedure of provisioning new services and products through eliciting preferences via the methods developed under (a) and (b).This Fellowship provides a unique opportunity to bring together artificial intelligence techniques and social science to tackle problems that are faced by a range of businesses and organisations in dealing with clients and customers and attempting to elicit preferences and needs through behaviours and interactions. We will be working closely with our industry partners in this project, British Telecom (BT) and the Essex County Council (ECC), to investigate the issues and challenges of eliciting and understanding preferences as being faced in their own contexts to inform and shape the programme of work.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间:
期刊:
影响因子: --
作者: [Amir Jahangiri]
通讯作者: Amir Jahangiri
Discovering Social and Individual Preferences with Feature Based Inverse Reinforcement Learning
通过基于特征的逆强化学习发现社会和个人偏好
DOI: --
发表时间: 2021
期刊:
影响因子: --
作者: [Amir Jahangiri]
通讯作者: Amir Jahangiri
海外基金