Welcome to D
SIGMOD 2005
PODS 2005
SIGMOD-RECOR
CIDR 2005
CIKM 2005
COMAD 2005
CVDB 2005
DaMoN 2005
Data Enginee
DEBS05
DMSN 2005
DOLAP 2005
GIR 2005
GIS 2005
Hypertext 20
ICDE 2005
ICDM 2005
IHIS 2005
IQIS 2005
JCDL 2005
KRAS 2005
MDM 2005
MIR 2005
MobiDE 2005
P2PIR 2005
RIDE 2005
<<< = RIDE'05 Pape>>>
SBBD 2005
SIGIR 2005
SIGIR-FORUM
SIGKDD 2005
SIGKDD-EXP
SSDBM 2005
TIME 2005
TKDE 2005
TODS 2005
VLDB 2005
VLDBJ 2005
WebDB 2005
WIDM 2005

Using Probabilistic Latent Semantic Analysis for Web Page Grouping


Guandong Xu, Yanchun Zhang, and Xiaofang Zhou

  View Paper (PDF)  

Return to Applications of Stream Data Mining


Abstract

The locality of Web pages within a Web site is initially determined by the designer's expectation. Web usage mining can discover the patterns in the navigational behaviour of Web visitors, in turn, improve Web site functionality and service designing by considering users' actual opinion. Conventional Web page clustering technique is often utilized to reveal the functional similarity of Web pages. However, high-dimensional computation problem will be incurred due to taking user transaction as dimension. In this paper, we propose a new Web page grouping approach based on a probabilistic latent semantic analysis (PLSA) model. An iterative algorithm based on maximum likelihood principle is employed to overcome the aforementioned computational shortcoming. The Web pages are classified into various groups according to user access patterns. Meanwhile, the semantic latent factors or tasks are characterized by extracting the content of "dominant" pages related to the factors. We demonstrate the effectiveness of our approach by conducting experiments on real world data sets.


©2006 Association for Computing Machinery