Welcome to DiSC 2003
SIGMOD 2002
PODS 2002
 SIGMOD RECORD 2002
 ADBIS 2002
CIKM 2002
CoopIS 2002
 EDBT 2002
 ER 2002
Data Engineering Bul
DEXA_EC-WEB 2002
DMKD 2002
 DPDJ 2002
HYPERTEXT 2002
ICDE 2002
ICDM 2002
<<< = ICDM'02 papers>>>
JCDL 2002
KDD 2002
 KDD_EXPLORATIONS 20
KRDB 2002
MDM 2002
MIS 2002
RIDE 2002
SBBD 2002
 SIGIR 2002
 SIGIR FORUM 2002
SSDBM 2002
TODS 2002
TIME 2002
VLDB 2002
VLDBJ 2002

Estimating the number of segments in time series data using permutation tests


Kari Vasko and Hannu Toivonen

  View Paper (PDF)  

Return to Main-Track Regular Papers


Abstract

Segmentation is a popular technique for discovering structure in time series data. We address the largely open problem of estimating the number of segments that can be reliably discovered. We introduce a novel method for the problem, called Pete. Pete is based on permutation testing. The problem is an instance of model (dimension) selection. The proposed method analyzes the possible overfit of a model to the available data rather than uses a term for penalizing model complexity. In this respect the approach is more similar to cross-validation than regulariza-tion based techniques (e.g., AIC, BIC, MDL, MML). Further, the method produces a p value for each increase in the number of segments. This gives the user an overview of the statistical significance of the segmentations. We evaluate the performance of the proposed method using both synthetic and real time series data. The experiments show that permutation testing gives realistic results about the number of reliably identifiable segments and that it compares favorably with the Monte Carlo cross-validation (MCCV) and commonly used BIC criteria.


DiSC'03 © 2003 Association for Computing Machinery