通过平均的分布式线性回归

Distributed linear regression by averaging

Annals of Statistics · 2021

被引 47

ABS 4*

Edgar Dobriban
Yue Sheng

中文导读

研究了在数据并行下，通过加权平均参数进行分布式线性回归的性能损失，发现估计误差和置信区间长度显著增加，而预测误差增加较少。

Abstract

Distributed statistical learning problems arise commonly when dealing with large datasets. In this setup, datasets are partitioned over machines, which compute locally, and communicate short messages. Communication is often the bottleneck. In this paper, we study one-step and iterative weighted parameter averaging in statistical linear models under data parallelism. We do linear regression on each machine, send the results to a central server and take a weighted average of the parameters. Optionally, we iterate, sending back the weighted average and doing local ridge regressions centered at it. How does this work compared to doing linear regression on the full data? Here, we study the performance loss in estimation and test error, and confidence interval length in high dimensions, where the number of parameters is comparable to the training data size. We find the performance loss in one-step weighted averaging, and also give results for iterative averaging. We also find that different problems are affected differently by the distributed framework. Estimation error and confidence interval length increases a lot, while prediction error increases much less. We rely on recent results from random matrix theory, where we develop a new calculus of deterministic equivalents as a tool of broader interest.

分布式统计学习线性回归高维统计随机矩阵理论

阅读原文 ↗