基于半参数模型的动态定价策略优化

Policy Optimization Using Semiparametric Models for Dynamic Pricing

Journal of the American Statistical Association · 2022
被引 26
ABS 4

中文导读

研究了上下文动态定价问题,提出一种结合半参数估计与在线决策的策略,在温和条件下实现了接近理论下界的遗憾上界,适用于市场噪声分布未知的场景。

Abstract

In this article, we study the contextual dynamic pricing problem where the market value of a product is linear in its observed features plus some market noise. Products are sold one at a time, and only a binary response indicating success or failure of a sale is observed. Our model setting is similar to the work by? except that we expand the demand curve to a semiparametric model and learn dynamically both parametric and nonparametric components. We propose a dynamic statistical learning and decision making policy that minimizes regret (maximizes revenue) by combining semiparametric estimation for a generalized linear model with unknown link and online decision making. Under mild conditions, for a market noise cdf F(·) with mth order derivative ( m≥2), our policy achieves a regret upper bound of O˜d(T2m+14m−1), where T is the time horizon and O˜d is the order hiding logarithmic terms and the feature dimension d. The upper bound is further reduced to O˜d(T) if F is super smooth. These upper bounds are close to Ω(T), the lower bound where F belongs to a parametric class. We further generalize these results to the case with dynamic dependent product features under the strong mixing condition. Supplementary materials for this article are available online.

动态定价半参数模型在线学习遗憾最小化计量经济学