通过代表性区域选择进行数据探索:公理与收敛性

Data Exploration by Representative Region Selection: Axioms and Convergence

Mathematics of Operations Research · 2021
被引 1
ABS 3

中文导读

提出一种新的无监督学习问题,通过选取少量代表性区域来近似大数据集,帮助实践者探索数据,不依赖聚类结构,并给出公理、质量函数和收敛性结果。

Abstract

We present a new type of unsupervised learning problem in which we find a small set of representative regions that approximates a larger data set. These regions may be presented to a practitioner along with additional information in order to help the practitioner explore the data set. An advantage of this approach is that it does not rely on cluster structure of the data. We formally define this problem, and we present axioms that should be satisfied by functions that measure the quality of representatives. We provide a quality function that satisfies all of these axioms. Using this quality function, we formulate two optimization problems for finding representatives. We provide convergence results for a general class of methods, and we show that these results apply to several specific methods, including methods derived from the solution of the optimization problems formulated in this paper. We provide an example of how representative regions may be used to explore a data set.

数据挖掘无监督学习优化人工智能