Efficient Data Stream Clustering with Sliding Windows based on Locality-Sensitive Hashing

Youn, Jonghem; Shim, Junho; Lee, Sang-Goo

doi:10.1109/ACCESS.2018.2877138

상세 보기

Efficient Data Stream Clustering with Sliding Windows based on Locality-Sensitive Hashing

Youn, Jonghem;
Shim, Junho;
Lee, Sang-Goo

Citations

WEB OF SCIENCE

17

Citations

SCOPUS

30

초록

Data stream clustering over sliding windows generates clusters as the window moves. However, iterative clustering using all data in a window is highly inefficient in terms of memory use and computational load. In this paper, we improve data stream clustering over sliding windows using sliding window aggregation and nearest neighbor search techniques. Our algorithm constructs and maintains temporal group features as a summary of the window using the sliding window aggregation technique. In order to maintain a constant size for the summary, the algorithm reduces the size of the summary by joining the nearest neighbor. We exploit locality-sensitive hashing for rapid nearest neighbor searching. In addition, we also suggest a re-clustering policy that determines whether to append a new summary to pre-existing clusters or to perform clustering on the whole summary. We conduct experiments on real-world and synthetic datasets in order to demonstrate that our algorithm can significantly improve continuous clustering on data streams with sliding windows.

키워드

Data stream; k-means clustering; locality-sensitive hashing; sliding window; EVOLVING DATA STREAMS; AFFINITY PROPAGATION

제목: Efficient Data Stream Clustering with Sliding Windows based on Locality-Sensitive Hashing

저자: Youn, Jonghem; Shim, Junho; Lee, Sang-Goo

DOI: 10.1109/ACCESS.2018.2877138

발행일: 2018-10

유형: Article

저널명: IEEE Access

권: 6

페이지: 63757 ~ 63776

ScholarWorks@숙명여자대학교

상세 보기

초록

키워드