Welcome to D
SIGMOD'00
 = SIGMOD'00 We
 = Plenary Talk
<<< = SIGMOD'00 Pa>>>
PODS'00
SIGMOD Recor
CIKM 2000/CI
COMAD 2000
Data Enginee
DL 2000
DPDJ
EDBT 2000
Hypertext 20
ICDE 2000
KDD 2000
KDD Explorat
KRDB 2000
SBBD 2000
SIGIR 2000
SIGIR Forum
SSDBM 2000
TODS
VLDB'00
VLDBJ

Finding Replicated Web Collections


Junghoo Cho, Narayanan Shivakumar, and Hector Garcia-Molina

  View Paper (PDF)  

Return to Research Sessions


Abstract

Many web documents (such as JAVA FAQs) are being replicated on the Internet. Often entire document collections (such as hyperlinked Linux manuals) are being replicated many times. In this paper, we make the case for identifying replicated documents and collections to improve web crawlers, archivers, and ranking functions used in search engines. The paper describes how to efficiently identify replicated documents and hyperlinked document collections. The challenge is to identify these replicas from an input data set of several tens of millions of web pages and several hundreds of gigabytes of textual data. We also present two real-life case studies where we used replication information to improve a crawler and a search engine. We report these results for a data set of 25 million web pages (about 150 gigabytes of HTML data) crawled from the web.


References


Note: References link to DBLP on the Web.

[1]
...
[2]
Krishna Bharat , Andrei Z. Broder : Mirror, Mirror on the Web: A Study of Host Pairs with Replicated Content. WWW8 / Computer Networks 31(11-16) : 1579-1590(1999)
[3]
Google. http://www.google.com
[4]
...
[5]
Andrei Z. Broder , Steven C. Glassman , Mark S. Manasse , Geoffrey Zweig : Syntactic Clustering of the Web. WWW6 / Computer Networks 29(8-13) : 1157-1166(1997)
[6]
Thomas H. Cormen , Charles E. Leiserson , Ronald L. Rivest : Introduction to Algorithms. The MIT Press and McGraw-Hill Book Company 1989, ISBN 0-262-03141-8 and 0-07-013143-0
[7]
Min Fang , Narayanan Shivakumar , Hector Garcia-Molina , Rajeev Motwani , Jeffrey D. Ullman : Computing Iceberg Queries Efficiently. VLDB 1998 : 299-310
[8]
...
[9]
Mike Perkowitz , Oren Etzioni : Adaptive Web Sites: Automatically Synthesizing Web Pages. AAAI/IAAI 1998 : 727-732
[10]
James E. Pitkow , Peter Pirolli : Life, Death, and Lawfulness on the Electronic Frontier. CHI 1997 : 383-390
[11]
Gerard Salton , Michael McGill : Introduction to Modern Information Retrieval. McGraw-Hill Book Company 1984, ISBN 0-07-054484-0
[12]
Narayanan Shivakumar , Hector Garcia-Molina : SCAM: A Copy Detection Mechanism for Digital Documents. DL 1995 : 0-
[13]
Narayanan Shivakumar , Hector Garcia-Molina : Building a Scalable and Accurate Copy Detection Mechanism. Digital Libraries 1996 : 160-168

BIBTEX


@inproceedings{DBLP:conf/sigmod/ChoSG00,
  author    = {Junghoo Cho and
                Narayanan Shivakumar and
                Hector Garcia-Molina},
   editor    = {Weidong Chen and
                Jeffrey F. Naughton and
                Philip A. Bernstein},
   title     = {Finding Replicated Web Collections},
   booktitle = {Proceedings of the 2000 ACM SIGMOD International Conference on
                Management of Data, May 16-18, 2000, Dallas, Texas, USA},
   journal   = {SIGMOD Record},
   publisher = {ACM},
   volume    = {29},
   number    = {2},
   year      = {2000},
   isbn      = {1-58113-218-2},
   pages     = {355-366},
   crossref  = {DBLP:conf/sigmod/2000},
   bibsource = {DBLP, http://dblp.uni-trier.de} } },




DiSC'01 Copyright ©2002 ACM Inc.