面向删失生存数据建模的具有普适适应性的多校准方法

Multicalibration for modelling censored survival data with universal adaptability

Biometrika · 2025
被引 1
ABS 4

中文导读

提出一种针对删失生存数据的黑箱后处理提升算法,通过多校准确保在目标域中预测生存概率和受限平均生存时间的准确性与公平性,理论证明其与逆概率加权估计相当或更优。

Abstract

Summary Traditional statistical and machine learning methods typically assume that the training and test data follow the same distribution. However, this assumption is frequently violated in real-world applications, where the training data in the source domain may underrepresent specific subpopulations in the test data of the target domain. This article addresses target-independent learning under covariate shift, focusing on multicalibration for survival probability and restricted mean survival time. A black-box post-processing boosting algorithm specifically designed for censored survival data is introduced. By leveraging pseudo-observations, our method produces a multicalibrated predictor that is competitive with inverse propensity score weighting in predicting the survival outcome in an unlabelled target domain, ensuring, not only overall accuracy, but also fairness across diverse subpopulations. Our theoretical analysis of pseudo-observations builds upon the functional delta method and the $ p $-variational norm. The algorithm’s sample complexity, convergence properties and multicalibration guarantees for post-processed predictors are provided. The results establish a fundamental connection between multicalibration and universal adaptability, demonstrating that our calibrated function is comparable to, or outperforms, the inverse propensity score weighting estimator. Extensive numerical simulations and a real-world case study on cardiovascular disease risk prediction using two large prospective cohort studies validate the effectiveness of our approach.

生存分析机器学习因果推断公平性