A polythetic clustering process and cluster validity indexes for histogram-valued objects
- Jaejik Kim(corresponding author),
- L. Billard
- ,
- University of Georgia
Abstract
Clustering is an explanatory procedure which helps to understand data with complex structure and multivariate relationships, and is a very useful method to extract knowledge and information especially from large datasets. When such datasets are aggregated into categories (as driven by scientific questions underlying the analysis), the resulting observations will perforce be expressed as so-called symbolic data (though symbolic data can occur "naturally" in any sized datasets). The focus of this work is to provide a divisive polythetic algorithm to establish clusters for p-dimensional histogram-valued data. In addition, two cluster validity indexes for use in establishing the optimal number of clusters are also developed. Finally, the proposed procedure is applied to a large forestry cover type dataset.
Publication Information
Output type
Original language
English (US)Pages from-to (Number of pages)
Pages 2250-2262 (13 pages)Journal (Volume, Issue Number)
Computational Statistics and Data Analysis (Volume 55, Issue 7)Publication milestones
- Published - 07/01/2011
Publication status
ISSN
0167-9473Publication IDs
- Scopus: 79953671001
