Analysis of Olympic Heptathlon Data
描述了一种增量聚类算法,应用于1992年奥运会七项全能数据,对顶尖运动员分组,并与经典聚类方法比较,可识别决定优秀运动员的关键项目。
Abstract An incremental clustering algorithm is described and applied to 1992 Olympic heptathlon data to produce characterizations of groups of the leading athletes. Results are compared with those obtained by classical clustering techniques. The incremental technique is of order n log n, where n is the number of observational units and can be combined with Classification and Regression Trees (CART) to obtain descriptions of the clusters. Brief reference is made to other possible ways of analyzing the data, including the use of correspondence analysis. The analysis can be used to show which events are most critical in determining the better class of heptathlete.