Model calibration and evaluation via optimal subsampling using electronic health record data
针对电子健康档案数据中结局标签缺失的问题,提出一种两步最优抽样设计,通过最小化预测准确性指标估计的渐近方差,高效获取有信息量的标签子集,实现模型的无偏验证。
Abstract A common challenge for validating risk prediction models using electronic health record (EHR) data is that labels for the predicted outcome are not directly available. Towards efficient and unbiased model validation, we study optimal sampling designs for efficiently labelling an informative subset of patients in an EHR cohort. Given a pre-specified number of outcome labels, our design aims to minimize the asymptotic variance of an improved inverse probability weighted (‘I-IPW’) estimator for predictive accuracy metrics. Implementation of the sampling requires accurate risk estimates and the predictive accuracy metric of interest. We therefore propose to implement sampling in two steps. First a portion of the target number of labels is acquired by applying entropy sampling to a random subset of the cohort. These initial labels are used to calibrate risk estimates and obtain an initial estimate of the predictive accuracy metric, which are used to inform optimal sampling of the remaining target number of labels. The final estimate of the predictive accuracy metrics is obtained by applying the I-IPW estimator to the cohort and all acquired labels pooled together. Results from simulation studies and application to a real EHR dataset indicate superior efficiency of the proposed sampling design and I-IPW estimator.