Performance Guarantees for Policy Learning.
Performance Guarantees for Policy Learning.
复制标题
DOI:
10.1214/19-aihp1034
复制
发表时间:
2020-08
期刊:
影响因子:
--
通讯作者:
Chambaz A
中科院分区:
文献类型:
--
作者:
Luedtke A;Chambaz A
This article gives performance guarantees for the regret decay in optimal policy estimation. We give a margin-free result showing that the regret decay for estimating a within-class optimal policy is second-order for empirical risk minimizers over Donsker classes when the data are generated from a fixed data distribution that does not change with sample size, with regret decaying at a faster rate than the standard error of an efficient estimator of the value of an optimal policy. We also present a result giving guarantees on the regret decay of policy estimators for the case that the policy falls within a restricted class and the data are generated from local perturbations of a fixed distribution, where this guarantee is uniform in the direction of the local perturbation. Finally, we give a result from the classification literature that shows that faster regret decay is possible via plug-in estimation provided a margin condition holds. Three examples are considered. In these examples, the regret is expressed in terms of either the mean value or the median value, and the number of possible actions is either two or finitely many.
登录
查看更多内容
影响因子:
4.5
作者:
Qian M;Murphy SA
通讯作者:
Murphy SA
影响因子:
6.1
作者:
Hirano, Keisuke;Porter, Jack R.
通讯作者:
Porter, Jack R.
影响因子:
4.5
作者:
Chambaz, Antoine;Zheng, Wenjing;van der Laan, Mark J.
通讯作者:
van der Laan, Mark J.
影响因子:
4.5
作者:
Audibert, Jean-Yves;Tsybakov, Alexandre B.
通讯作者:
Tsybakov, Alexandre B.
DOI:
10.1515/1557-4679.1423
发表时间:
2012-07-20
期刊:
The international journal of biostatistics
影响因子:
--
作者:
Rubin, Daniel B;van der Laan, Mark J
通讯作者:
van der Laan, Mark J