Anytime-Valid Inference in Linear Models with Applications to Regression-Adjusted Causal Inference
本文为线性模型开发了任意时点有效的推断理论,引入序贯版本的经典检验和置信集,能在所有样本量下控制第一类错误,并应用于在线A/B测试的回归调整因果推断。
Linear models are foundational tools in statistics and ubiquitous across the applied sciences. However, conventional statistical tests, such as t-tests and F-tests, are only valid at fixed sample sizes, making them unsuitable for sequential settings such as online A/B testing. We develop an anytime-valid theory of inference for the linear model, introducing sequential analogues of classical tests and confidence sets that provide Type-I error control and coverage guarantees uniformly over all sample sizes. Our construction is based on likelihood ratios of invariantly sufficient statistics, yielding simple closed-form expressions of ordinary least squares estimators and standard errors. The resulting tests are optimal in the GROW/REGROW sense for both frequentist and Bayesian alternative hypotheses. We then relax the linear model assumptions to provide heteroskedasticity-robust asymptotic sequential tests and confidence sequences, which enable sequential regression-adjusted inference for causal estimands in randomized controlled experiments. This formally allows experiments to be continuously monitored for significance, stopped early, and safeguarded against statistical malpractices in data collection. We demonstrate the practical utility of our approach through simulations and applications to real A/B test data from Netflix.