2011 MethodsforMinFreqItemsinDataStrAnOver
- (Hongyan &al, 2011) ⇒ Hongyan Liu, Yuan Lin and Jiawei Han. (2011). “Methods for Mining Frequent Items in Data Streams: An Overview.” In: International Journal on Knowledge and Information Systems (KAIS), (26:1). doi:10.1007/s10115-009-0267-2
Subject Headings: Data mining, Data stream, Mining methods and algorithms, Frequent items.
Notes
Cited By
Quotes
Abstract
In many real-world applications, information such as web click data, stock ticker data, sensor network data, phone call records, and traffic monitoring data appear in the form of data streams. Online monitoring of data streams has emerged as an important research undertaking. Estimating the frequency of the items on these streams is an important aggregation and summary technique for both stream mining and data management systems with a broad range of applications. This paper reviews the state-of-the-art progress on methods of identifying frequent items from data streams. It describes different kinds of models for frequent items mining task. For general models such as cash register and Turnstile, we classify existing algorithms into sampling-based, counting-based, and hashing-based categories. The processing techniques and data synopsis structure of each algorithm are described and compared by evaluation measures. Accordingly, as an extension of the general data stream model, four more specific models including time-sensitive model, distributed model, hierarchical and multi-dimensional model, and skewed data model are introduced. The characteristics and limitations of the algorithms of each model are presented, and open issues waiting for study and improvement are discussed.
References
,
Author | volume | Date Value | title | type | journal | titleUrl | doi | note | year | |
---|---|---|---|---|---|---|---|---|---|---|
2011 MethodsforMinFreqItemsinDataStrAnOver | Hongyan Liu Yuan Lin Jiawei Han | Methods for Mining Frequent Items in Data Streams: An Overview | International Journal on Knowledge and Information System | http://www.springerlink.com/content/0797w54j30k28257/ | 2011 |