Welcome to D
SIGMOD'00
 = SIGMOD'00 We
 = Plenary Talk
<<< = SIGMOD'00 Pa>>>
PODS'00
SIGMOD Recor
CIKM 2000/CI
COMAD 2000
Data Enginee
DL 2000
DPDJ
EDBT 2000
Hypertext 20
ICDE 2000
KDD 2000
KDD Explorat
KRDB 2000
SBBD 2000
SIGIR 2000
SIGIR Forum
SSDBM 2000
TODS
VLDB'00
VLDBJ

Turbo-charging Vertical Mining of Large Databases


Pradeep Shenoy, Jayant R. Haritsa, S. Sudarshan, Gaurav Bhalotia, Mayank Bawa, and Devavrat Shah

  View Paper (PDF)  

Return to Research Sessions


Abstract

In a vertical representation of a market-basket database, each item is associated with a column of values representing the transactions in which it is present. The association-rule mining algorithms that have been recently proposed for this representation show performance improvements over their classical horizontal counterparts, but are either efficient only for certain database sizes, or assume particular characteristics of the database contents, or are applicable only to specific kinds of database schemas. We present here a new vertical mining algorithm called VIPER, which is general-purpose, making no special requirements of the underlying database. VIPER stores data in compressed bit-vectors called "snakes" and integrates a number of novel optimizations for efficient snake generation, intersection, counting and storage. We analyze the performance of VIPER for a range of synthetic database workloads. Our experimental results indicate significant performance gains, especially for large databases, over previously proposed vertical and horizontal mining algorithms. In fact, there are even workload regions where VIPER outperforms an optimal, but practically infeasible, horizontal mining algorithm.


References


Note: References link to DBLP on the Web.

[1]
Rakesh Agrawal , Tomasz Imielinski , Arun N. Swami : Mining Association Rules between Sets of Items in Large Databases. SIGMOD Conference 1993 : 207-216
[2]
Rakesh Agrawal , Ramakrishnan Srikant : Fast Algorithms for Mining Association Rules in Large Databases. VLDB 1994 : 487-499
[3]
Brian Dunkel , Nandit Soparkar : Data Organization and Access for Efficient Data Mining. ICDE 1999 : 522-529
[4]
...
[5]
...
[6]
Marcel Holsheimer , Martin L. Kersten , Heikki Mannila , Hannu Toivonen : A Perspective on Databases and Data Mining. KDD 1995 : 150-155
[7]
Ashoka Savasere , Edward Omiecinski , Shamkant B. Navathe : An Efficient Algorithm for Mining Association Rules in Large Databases. VLDB 1995 : 432-444
[8]
...
[9]
Show-Jane Yen , Arbee L. P. Chen : An Efficient Approach to Discovering Knowledge from Large Databases. PDIS 1996 : 8-18
[10]
...
[11]
Mohammed Javeed Zaki , Srinivasan Parthasarathy , Mitsunori Ogihara , Wei Li : New Algorithms for Fast Discovery of Association Rules. KDD 1997 : 283-286

BIBTEX


@inproceedings{DBLP:conf/sigmod/ShenoyHSBBS00,
  author    = {Pradeep Shenoy and
                Jayant R. Haritsa and
                S. Sudarshan and
                Gaurav Bhalotia and
                Mayank Bawa and
                Devavrat Shah},
   editor    = {Weidong Chen and
                Jeffrey F. Naughton and
                Philip A. Bernstein},
   title     = {Turbo-charging Vertical Mining of Large Databases},
   booktitle = {Proceedings of the 2000 ACM SIGMOD International Conference on
                Management of Data, May 16-18, 2000, Dallas, Texas, USA},
   journal   = {SIGMOD Record},
   publisher = {ACM},
   volume    = {29},
   number    = {2},
   year      = {2000},
   isbn      = {1-58113-218-2},
   pages     = {22-33},
   crossref  = {DBLP:conf/sigmod/2000},
   bibsource = {DBLP, http://dblp.uni-trier.de} } },




DiSC'01 Copyright ©2002 ACM Inc.