U-Statistic Reduction: Higher-Order Accurate Risk Control and Statistical-Computational Trade-Off
本文针对U统计量计算慢的问题,提出了首个能保证高阶精确风险控制的不完全U统计量推断方法,揭示了风险控制精度与计算速度之间的权衡,并通过数值和实际数据验证了方法的有效性。
U-statistics play central roles in many statistical learning tools but face the haunting issue of scalability. Despite extensive research on accelerating computation by U-statistic reduction, existing results almost exclusively focused on power analysis. Little work addresses risk control accuracy, which requires distinct and much more challenging techniques. In this paper, we establish the first statistical inference procedure with provably higher-order accurate risk control for incomplete U-statistics. The sharpness of our new result enables us to reveal how risk control accuracy also trades off with speed, for the first time in literature, which complements the well-known variance-speed trade-off. Our general framework converts the challenging and case-by-case analysis for many different designs into a surprisingly principled and routine computation. We conducted comprehensive numerical studies and observed results that validate our theory’s sharpness. Our method also demonstrates effectiveness on real-world data applications.