Learning Performance of Weighted Distributed Learning With Support Vector Machines
针对大数据中噪声样本影响算法性能的问题,引入马尔可夫采样和权重,研究了加权分布式支持向量机的泛化误差和最优收敛速度,并提出了基于马尔可夫采样的新算法,在基准数据集上表现更优且耗时更少。
The divide-and-conquer strategy is a very effective method of dealing with big data. Noisy samples in big data usually have a great impact on algorithmic performance. In this article, we introduce Markov sampling and different weights for distributed learning with the classical support vector machine (cSVM). We first estimate the generalization error of weighted distributed cSVM algorithm with uniformly ergodic Markov chain (u.e.M.c.) samples and obtain its optimal convergence rate. As applications, we obtain the generalization bounds of weighted distributed cSVM with strong mixing observations and independent and identically distributed (i.i.d.) samples, respectively. We also propose a novel weighted distributed cSVM based on Markov sampling (DM-cSVM). The numerical studies of benchmark datasets show that the DM-cSVM algorithm not only has better performance but also has less total time of sampling and training compared to other distributed algorithms.