2011 INCONCOInterpretableClusteringo
- (Plant & Böhm, 2011) ⇒ Claudia Plant, and Christian Böhm. (2011). “INCONCO: Interpretable Clustering of Numerical and Categorical Objects.” In: Proceedings of the 17th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD-2011) Journal. ISBN:978-1-4503-0813-7 doi:10.1145/2020408.2020584
Subject Headings:
Notes
Cited By
- http://scholar.google.com/scholar?q=%222011%22+INCONCO%3A+Interpretable+Clustering+of+Numerical+and+Categorical+Objects
- http://dl.acm.org/citation.cfm?id=2020408.2020584&preflayout=flat#citedby
Quotes
Author Keywords
Abstract
The integrative mining of heterogeneous data and the interpretability of the data mining result are two of the most important challenges of today's data mining. It is commonly agreed in the community that, particularly in the research area of clustering, both challenges have not yet received the due attention. Only few approaches for clustering of objects with mixed-type attributes exist and those few approaches do not consider cluster-specific dependencies between numerical and categorical attributes. Likewise, only a few clustering papers address the problem of interpretability: to explain why a certain set of objects have been grouped into a cluster and what a particular cluster distinguishes from another. In this paper, we approach both challenges by constructing a relationship to the concept of data compression using the Minimum Description Length principle: a detected cluster structure is the better the more efficient it can be exploited for data compression. Following this idea, we can learn, during the run of a clustering algorithm, the optimal trade-off for attribute weights and distinguish relevant attribute dependencies from coincidental ones. We extend the efficient Cholesky decomposition to model dependencies in heterogeneous data and to ensure interpretability. Our proposed algorithm, INCONCO, successfully finds clusters in mixed type data sets, identifies the relevant attribute dependencies, and explains them using linear models and case-by-case analysis. Thereby, it outperforms existing approaches in effectiveness, as our extensive experimental evaluation demonstrates.
References
;
Author | volume | Date Value | title | type | journal | titleUrl | doi | note | year | |
---|---|---|---|---|---|---|---|---|---|---|
2011 INCONCOInterpretableClusteringo | Christian Böhm Claudia Plant | INCONCO: Interpretable Clustering of Numerical and Categorical Objects | 10.1145/2020408.2020584 | 2011 |