Welcome to D
SIGMOD 2004
PODS 2004
SIGMOD RECOR
CIKM 2004
DASFAA 2004
DBPL 2003
DE-BULLETIN
DEBS 2004
DMKD 2004
DMSN 2004
DOLAP 2004
DPDJ 2004
EDBT 2004
ER 2003
GIS 2004
HDP 2004
HYPERTEXT 20
ICDE 2004
ICDT 2003
JCDL 2004
MDM
MIR 2004
MIS 2004
MMDB 2004
MOBIDE 2003
RIDE 2004
SBBD 2003
SIGIR FORUM
SIGIR 2004
<<< = SIGIR'04 Pap>>>
SIGKDD EXPLO
SIGKDD 2004
SSDBM 2004
SSTD 2003
TIME 2004
TODS 2004
VLDB 2004
VLDB Journal
WEBDB 2004
WIDM 2004
XIME-P 2004
Footer

Constructing a text corpus for inexact duplicate detection


Jack G. Conrad and Cindy P. Schriber

  View Paper (PDF)  

Return to Posters


Abstract

As online document collections continue to expand, both on the Web and in proprietary environments, the need for duplicate detection becomes more critical. The goal of this work is to facilitate (a) investigations into the phenomenon of near duplicates and (b) algorithmic approaches to minimizing its negative effect on search results. Harnessing the expertise of both client-users and professional searchers, we establish principled methods to generate a test collection for identifying and handling inexact duplicate documents.

BIBTEX


@inproceedings{1009131,   author = {Jack G. Conrad and Cindy P. Schriber},
  title = {Constructing a text corpus for inexact duplicate detection},
  booktitle = {SIGIR '04: Proceedings of the 27th annual international conference on Research and development in information retrieval},
  year = {2004},
  isbn = {1-58113-881-4},
  pages = {582--583},
  location = {Sheffield, United Kingdom},
  doi = {http://doi.acm.org/10.1145/1008992.1009131},
  publisher = {ACM Press},
  
}



©2005 Association for Computing Machinery