Multi-Modal Dynamic Pricing

Multi-Modal Dynamic Pricing
复制标题

多模式动态定价

DOI:
--
复制
发表时间:
2019
期刊:
Management Sciences
影响因子:
--
通讯作者:
D. Simchi
D. Simchi
中科院分区:
--
文献类型:
--
作者:
Yining Wang;Boxiao Chen;D. Simchi

文献摘要

被引文献

相似文献

研究了具有需求学习的单产品动态定价问题.候选价格属于一个广泛的价格区间,建模的需求函数是非参数的性质,只施加平滑规律性条件。我们的模型的一个重要方面是预期奖励函数是非凹的,实际上是多峰的可能性,这导致了许多概念和技术上的挑战。我们所提出的算法的灵感来自于多臂土匪的上置信界算法和线性上下文土匪所产生的乐观面对不确定性原则。多臂的强盗制定产生于一个未知的连续需求函数的局部箱近似,然后应用线性上下文的强盗制定,以获得更准确的局部多项式逼近在每个箱。通过严格的遗憾分析,我们表明,我们提出的算法实现了最佳的最坏情况下的遗憾在广泛的光滑函数类。更具体地说,对于k次光滑函数和T个销售周期,我们提出的算法的遗憾是[公式:见正文],这是通过信息理论下限的发展被证明是最优的。我们还表明,在特殊情况下,如强凹或无限光滑的奖励函数,我们的算法实现了[公式:见文字]遗憾,匹配在以前的作品中建立的最佳遗憾。最后,我们提出的计算结果,验证了我们的方法在数值模拟的有效性。这篇论文被大数据分析师J.乔治·尚蒂库马尔接受。
We consider a single product dynamic pricing with demand learning. The candidate prices belong to a wide range of a price interval; the modeling of the demand functions is nonparametric in nature, imposing only smoothness regularity conditions. One important aspect of our model is the possibility of the expected reward function to be nonconcave and indeed multimodal, which leads to many conceptual and technical challenges. Our proposed algorithm is inspired by both the Upper-Confidence-Bound algorithm for multiarmed bandit and the Optimism-in-the-Face-of-Uncertainty principle arising from linear contextual bandits. The multiarmed bandit formulation arises from local-bin approximation of an unknown continuous demand function, and the linear contextual bandit formulation is then applied to obtain more accurate local polynomial approximators within each bin. Through rigorous regret analysis, we demonstrate that our proposed algorithm achieves optimal worst-case regret over a wide range of smooth function classes. More specifically, for k-times smooth functions and T selling periods, the regret of our proposed algorithm is [Formula: see text], which is shown to be optimal via the development of information theoretical lower bounds. We also show that in special cases, such as strongly concave or infinitely smooth reward functions, our algorithm achieves an [Formula: see text] regret, matching optimal regret established in previous works. Finally, we present computational results that verify the effectiveness of our method in numerical simulations. This paper was accepted by J. George Shanthikumar, big data analytics.