Welcome to D
SIGMOD 2003
PODS 2003
SIGMOD-RECOR
ADBIS
CIDR 2003
CIKM 2003
DASFAA 2003
Data Enginee
DEBS
DMKD 2003
DOLAP 2003
DPDJ 2003
ER
GIS 2003
Hypertext 20
ICDE 2003
ICDM 2003
ICDT 2003
JCDL 2003
KRDB 2003
MIR 2003
MIS 2003
MMDB 2003
RIDE 2003
SBBD 2003
SIGIR 2003
SIGIR-FORUM
SIGKDD 2003
SIGKDD-EXP
SSDBM 2003
TIME 2003
TODS
VLDB 2003
<<< = VLDB'03 Pape>>>
 = Plenary Talk
VLDB Journal
WIDM 2003

Robust Estimation With Sampling and Approximate Pre-Aggregation


Chris Jermaine

  View Paper (PDF)  

Return to Metadata & Sampling (Session C6)


Abstract

The majority of data reduction techniques for approxi- mate query processing (such as wavelets, histograms, kernels, and so on) are not usually applicable to cate- gorical data. There has been something of a disconnect between research in this area and the reality of data- base data; much recent research has focused on approximate query processing over ordered or numeri- cal attributes, but arguably the majority of database attributes are categorical: country, state, job_title, color, sex, department , and so on. This paper considers the problem of approximation of aggregate functions over categorical data, or mixed categorical/numerical data. We propose a method based upon random sampling, called Approximate Pre-Aggregation (APA). The biggest drawback of sampling for aggregate function estimating is the sen- sitivity of sampling to attribute value skew, and APA uses several techniques to overcome this sensitivity. The increase in accuracy using APA compared to "plain vanilla" sampling is dramatic. For SUM and AVG queries, the relative error for random sampling alone is more than 700% greater than for sampling with APA. Even if stratified sampling techniques are used, the error is still between 28% and 175% greater than for APA.


©2004 Association for Computing Machinery