Causal Bandits: Learning Good Interventions via Causal Inference
Causal Bandits: Learning Good Interventions via Causal Inference
复制标题
因果强盗:通过因果推理学习良好的干预措施
DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
Mark D. Reid
中科院分区:
文献类型:
--
作者:
Finnian Lattimore;Tor Lattimore;Mark D. Reid
We study the problem of using causal models to improve the rate at which good interventions can be learned online in a stochastic environment. Our formalism combines multi-arm bandits and causal inference to model a novel type of bandit feedback that is not exploited by existing approaches. We propose a new algorithm that exploits the causal feedback and prove a bound on its simple regret that is strictly better (in all quantities) than algorithms that do not use the additional causal information.