关于Sander Greenland受邀论文《分歧与决策P值:一个在理论上值得区分并在实践中保持的区别》的讨论

Discussion on the SJS invited paper by Sander Greenland Divergence vs. DecisionP$$ P $$‐values: A Distinction worth making in theory and keeping in Practice

Scandinavian Journal of Statistics · 2023
被引 0
ABS 3

中文导读

本文讨论P值在模型拟合检验中作为证据的适用性,指出拟合优度P值虽不违反似然原理,但将任何非零P值视为负面证据会导致悖论,并强调统计学家需主观设定阈值。

Abstract

It is shown in the cited paper written by Michael Lavine and in several others works that the p $$ p $$ -value of the test-statistics is not a consistent measure of evidence in the context of testing alternative hypothesis. As Sander Greenland points out, these should not be confused with the p $$ p $$ -values of realized goodness of-fit-test statistics. Goodness-of-fit tests are useful sanity checks in order to decide, (and there would be also a decision to be taken there !) whether our model and/or assumptions are compatible with the available data, or we need to take a step back and look for what could be wrong. Greenland argues that although using the p $$ p $$ -value in a goodness of fit test violates the likelihood principle, these p $$ p $$ -values do not suffer the inconsistencies of p $$ p $$ -values arising from testing alternative hypothesis, and can be used as a measure of evidence against the model, (following Popper, scientific theories and models can be only falsified and never verified), on an absolute scale in the [ 0 , 1 ] $$ \left[0,1\right] $$ range, where a zero left-sided p $$ p $$ -value corresponds to perfect fit, and every p $$ p $$ -value strictly greater than zero should be considered as negative evidence of some level. I find this peculiar, unless we would be dealing only with fully deterministic models. Fitting the data is not the same as evidence and a conscious statistician is aware that overfitting hints to a miss-calibrated model and leads to the lack of predictive power. Let's consider a simple and useless model, where all the data coordinates are i.i.d. Gaussian with zero mean and variance σ 2 $$ {\sigma}^2 $$ , regardless of any additional information we may have, as the covariates in a regression model. For a fixed data sample X = ( X 1 , … , X n ) ∈ ℝ n $$ X=\left({X}_1,\dots, {X}_n\right)\in {\mathbb{R}}^n $$ By considering a model with σ 2 $$ {\sigma}^2 $$ large enough the observed value of the χ 2 $$ {\chi}^2 $$ -statistics 1 n σ 2 ∑ i = 1 n X i 2 $$ \frac{1}{n{\sigma}^2}\sum \limits_{i=1}^n{X}_i^2 $$ would be negligible, while for n $$ n $$ -large it would be approximately standard Gaussian under the hypothesized model, with p $$ p $$ -value as close to zero as we want. If we stick to the logic of the article the evidence against a model claiming that the data contains only noise would be always negligible when the variance of the noise in the model is assumed to be large enough. We disagree with this nihilistic point of view, of course a p $$ p $$ -value of the goodness-of-fit test statistics too far in the right tail is evidence of deleterious overfitting and miscalibration. Considering two-sided p $$ p $$ -values would not solve the problem, any departure of the goodness-of-fit test statistics from its median would be then interpreted as negative evidence of some level. This would lead to the conclusion that almost all possible datasets carry some level of negative evidence against the model, a conclusion which I find paradoxical and not useful. In goodness-of-fit testing the p $$ p $$ -value of the test statistics is used to assess whether the realized data is typical or non-typical under the assumed model, and every statistician has to decide subjectively, by her personal judgement and professional experience, or by following some conventions, where to put the threshold between typical and nontypical in the p $$ p $$ -value scale, and only beyond that threshold she would consider the p $$ p $$ -value as evidence against the model.

统计学假设检验模型拟合P值