基于Spark的高速大数据流最近邻分类

Nearest Neighbor Classification for High-Speed Big Data Streams Using Spark

IEEE Transactions on Systems, Man, and Cybernetics: Systems · 2017
被引 82
ABS 3

中文导读

提出一种基于Spark的增量式分布式最近邻分类器,通过分布式度量空间排序和增量实例选择,高效处理高速大数据流,实验证明其有效性。

Abstract

Mining massive and high-speed data streams among the main contemporary challenges in machine learning. This calls for methods displaying a high computational efficacy, with ability to continuously update their structure and handle ever-arriving big number of instances. In this paper, we present a new incremental and distributed classifier based on the popular nearest neighbor algorithm, adapted to such a demanding scenario. This method, implemented in Apache Spark, includes a distributed metric-space ordering to perform faster searches. Additionally, we propose an efficient incremental instance selection method for massive data streams that continuously update and remove outdated examples from the case-base. This alleviates the high computational requirements of the original classifier, thus making it suitable for the considered problem. Experimental study conducted on a set of real-life massive data streams proves the usefulness of the proposed solution and shows that we are able to provide the first efficient nearest neighbor solution for high-speed big and streaming data.

数据流挖掘大数据机器学习分布式计算