Kolyan Ray 和 Botond Szabo 对 Fong、Holmes 和 Walker 的《鞅后验分布》讨论的贡献

Kolyan Ray and Botond Szabo's contribution to the Discussion of ‘Martingale Posterior Distributions’ by Fong, Holmes and Walker

Journal of the Royal Statistical Society. Series B: Statistical Methodology · 2023
被引 0
ABS 4

中文导读

本文通过一个简单例子说明,在似然函数不连续时,鞅后验的预测分布可能不一致,提醒使用该方法需谨慎,并呼吁发展新的一致性条件。

Abstract

We congratulate the authors for their thought-provoking article, which includes proposing a copula-based update for the predictive density and establishing its frequentist consistency under relatively mild assumptions. In this contribution, we further explore the frequentist properties of (Bayesian) predictive densities and illustrate through a simple example that, similarly to the posterior distribution, the predictive distribution can also be inconsistent. One therefore requires caution when using martingale posteriors, at least for the frequentist. Consider a modified version of Example 1 in the present paper, taken from Christensen (2009). Let Y1,…,Yn∼iidfθ⁠, where i.e. for rational parameter values θ, we replace the N(θ,1) Gaussian density by a Cauchy with location parameter θ. As in Example 1, we endow θ with a standard Gaussian prior, i.e. π(θ)=N(θ∣0,1)⁠. Since the likelihoods in our example and Example 1 are equal almost everywhere under the prior, one can show the corresponding posteriors are also identical, i.e. the posterior is N(θ∣θ¯n,σ¯n2) with θ¯=∑i=1nyi/(n+1) and σ¯n2=1/(n+1)⁠, see Christensen (2009). Similarly, the posterior predictive is p(y∣y1:n)=N(y∣θ¯n,σ¯n2+1)⁠. By Example 2 in Hahn et al. (2018), the predictive updates can thus be characterised via a Gaussian copula with correlation parameter ρn=(1+n)−1⁠. However, for any fixed θ∈Q⁠, which forms a dense set in the parameter space R⁠, the data is Cauchy. The posterior is thus inconsistent for any rational ‘true’ parameter θ∈Q⁠, and the predictive density differs substantially from the true Cauchy density, even as n→∞⁠. Following the notation of the paper, this procedure cannot consistently recover θ∞=θ(Y1:∞)=fθ⁠, even though this parameter is fully defined by the infinite observations Y1:∞⁠. Of course, this example does not contradict Theorem 7, as the assumption ‖f0/p0‖∞≤B does not hold when f0 is Cauchy and p0 is Gaussian. In this simple example, the discontinuous likelihood function causes the posterior inconsistency, which in turn implies inconsistency of the predictive distribution. The issue is that the above approach only works for parameter values in a set of prior probability one, namely R∖Q⁠. However, prior null sets can be very large if not judged from the prior perspective. This problem becomes more pronounced in nonparametric models, where the set of parameters over which the posterior is consistent can be topologically negligible compared to those where it is inconsistent, see the classical results (Diaconis & Freedman, 1986; Freedman & Diaconis, 1983). This in turn results in predictive distributions not resembling the true data generating distribution. One must therefore be careful with the choice of predictive distribution, justifying the approach as is done in Theorem 7. We second the authors’ view that it is of interest to derive new conditions, analogous to the Kullback–Leibler property in the classical Bayesian setting, which yield consistency and ideally minimax concentration rates. Cofounded by the European Union (ERC, BigBayesUQ, project number: 101041064). Views and opinions expressed are, however, those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Council. Neither the European Union nor the granting authority can be held responsible for them.

贝叶斯统计鞅后验分布预测密度频率学派一致性