Tamás P. Papp、Paul Fearnhead 和 Chris Sherlock 对‘关于机器学习的概率与统计方面的讨论会议’的讨论贡献

Tamás P. Papp, Paul Fearnhead, and Chris Sherlock’s contribution to the Discussion of ‘the Discussion Meeting on Probabilistic and statistical aspects of machine learning’

Journal of the Royal Statistical Society. Series B: Statistical Methodology · 2024
被引 0
ABS 4

中文导读

评论了Benton等人的论文,该论文将扩散生成模型推广到一般空间和噪声过程,但作者质疑其成功原因、参数估计可行性、KL散度在低维流形上的适用性,并询问参数化选择和反不对称约束的优化优势。

Abstract

We provide comments only on the read paper “From Denoising Diffusions to Denoising Markov Models” by Benton, Shi, De Bortoli, Deligiannidis and Doucet. We congratulate the authors on an important contribution to the exciting area of generative models: providing a simple framework for applying the idea of diffusion generative models to data defined on general spaces and for general noising processes. Diffusion models have shown remarkable empirical performance for several applications, most notably as a way of defining implicit models for images. However it is unclear, to us at least, why these methods are so successful. They typically involve fitting highly parameterized models for the score, and it is unclear why one would expect to be able to find good parameter estimates for such models. Furthermore, such models may be sufficiently flexible to learn the exact denoising process for the samples they are trained on—thereby collapsing the diffusion model to a categorical sampler from the training data. Could the authors provide insight into why these methods work? Is there intuition as to what types of data their denoising Markov models will work well for? Diffusion models are often used when the data are supported on a low-dimensional manifold of the state space. However, the Kullback-Leibler (KL) divergence which these models implicitly minimize can be infinite when the supports of the distributions are disjoint. Can the proposed framework extend to objectives which do not have this issue? Or is the use of KL beneficial, in that it actually helps force the model to sample from close to the manifold? The proposed framework parameterizes the generative process in terms of a function β(x;t)⁠, whose optimal value is the marginal density of the noised data at time t. However, in practice the authors use different parameterizations, mainly based on differences in logβ(x;t)⁠. Can the authors give more intuition as to why this is the best choice? And do they think this would be the case in general? Finally, both parameterizations of the continuous-time Markov chain model revolve around logr(x,y;t)=logβ(x;t)−logβ(y;t). As far as we can tell, the authors do not impose the anti-symmetry logr(x,y;t)=−logr(y,x;t) when fitting r, which seems like some form of relaxation of the optimization problem. Are there advantages in fitting a more general functional form? Even if you fit a general function r, you can apply a post-hoc anti-symmetric correction: With the denoising formulation, this does not require additional neural network evaluations. We found that this correction produced almost identical results for the MNIST inpainting example (see Table 1). Image quality metrics for MNIST 14×14 inpainting, with standard errors in brackets

生成模型扩散模型机器学习概率统计