Learning prospect theory value function and reference point of a sequential decision maker

Learning prospect theory value function and reference point of a sequential decision maker
复制标题

DOI:
10.1109/cdc.2017.8264531
复制
发表时间:
2017-12
期刊:
2017 IEEE 56th Annual Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
Kamil Nar;L. Ratliff;S. Sastry
Kamil Nar;L. Ratliff;S. Sastry
中科院分区:
其他
文献类型:
--
作者:
Kamil Nar;L. Ratliff;S. Sastry

文献摘要

被引文献

相似文献

给定一个决策问题,一个人的参考点决定了结果是被视为收益还是损失,并影响决策。在本文中,我们假设一个人被给予相同的决策问题重复,和人选择一个行动,以最大化她的价值函数,而她的参考点可能会随着时间的推移而改变。我们估计的价值函数和参考点的人从观察到的行动,通过构建一个隐马尔可夫模型,并使用期望最大化算法。然后,我们测试所提出的算法上的数据集的纽约市出租车司机。
Given a decision problem, the reference point of a person determines whether the outcomes are perceived as gain or loss and influences the decision. In this paper, we assume that a person is given the same decision problem repeatedly, and the person chooses an action to maximize her value function while her reference point could possibly change over time. We estimate the value function and the reference point of the person from the observed actions by constructing a hidden Markov model and using the expectation-maximization algorithm. Then we test the suggested algorithm on the data set of New York City taxi drivers.