Welcome to DiSC 2002
SIGMOD 2001
PODS 2001
 SIGMOD RECORD 2001
CIKM 2001
CoopIS 2001
DASFAA 2001
DASFAA 2000
DBPL 2001
Data Engineering Bul
DEXA_EC-WEB 2001
DMKD 2001
 DPDJ 2001
HYPERTEXT 2001
ICDE 2001
ICDM 2001
ICDT 2001
JCDL 2001
KDD 2001
 KDD_EXPLORATIONS 20
KRDB 2001
MDM 2001
MIR 2001
MIS 2001
RIDE 2001
SBBD 2001
 SIGIR 2001
 SIGIR FORUM 2001
SSDBM 2001
 = SSDBM'01 Website
<<< = SSDBM'01 Papers>>>
SSTD 2001
TODS 2001
TIME 2001
VLDB 2001
VLDBJ 2001

Clustering High dimensional Massive Scientific Datasets


Ekow J. Otoo, Arie Shoshani, and Seung-won Hwang

  View Paper (PDF)  

Return to Clustering Techniques


Abstract

Many scientific applications can benefit from efficient clustering algorithm of massively large high dimensional datasets. However most of he developed algorithms are impractical to use when the amount of data is very large. Given N objects each defined by an M dimensional feature vector, any clustering technique for handling very large datasets in high dimensional space should run in time O(N) at best, and O(N log N) in the worst case, using no more than O(NM) storage, for it to be practical. A parallelized version of the same algorithm should achieve a linear speed-up in processing time with increasing number of processors. We introduce a hybrid algorithm called HyCeltyc, as an approach for clustering massively large high dimensional datasets. HyCeltyc, which stands for Hybrid Cell Density Clustering method, combines a cell-density based algorithm with a hierarchical agglomerative method to identify clusters in linear time. The main steps of the algorithm involve sampling, dimensionality reduction and selection of significant features on which to cluster the data.


DiSC'02 © 2003 Association for Computing Machinery