Learning the two parameters of the Poisson–Dirichlet distribution with a forensic application
针对法医学中稀有特征匹配问题,研究了用两种采样方法学习双参数泊松-狄利克雷分布的后验分布,并证明全贝叶斯推断比现有近似方法更准确。
Abstract In forensic science, the rare type match problem arises when the matching characteristic from the suspect and the crime scene is not in the reference database; hence, it is difficult to evaluate the likelihood ratio that compares the defense and prosecution hypotheses. A recent solution consists of modeling the ordered population probabilities according to the two‐parameter Poisson–Dirichlet distribution, which is a well‐known Bayesian nonparametric prior, and plugging the maximum likelihood estimates of the parameters into the likelihood ratio. We demonstrate that this approximation produces a systematic bias that fully Bayesian inference avoids. Motivated by this forensic application, we consider the need to learn the posterior distribution of the parameters that governs the two‐parameter Poisson–Dirichlet using two sampling methods: Markov Chain Monte Carlo and approximate Bayesian computation. These methods are evaluated in terms of accuracy and efficiency. Finally, we compare the likelihood ratio that is obtained by our proposal with the existing solution using a database of Y‐chromosome haplotypes.