![]() ![]() ![]() |
![]() |
|
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Return to Posters As online document collections continue to expand, both on the Web and in proprietary environments, the need for duplicate detection becomes more critical. The goal of this work is to facilitate (a) investigations into the phenomenon of near duplicates and (b) algorithmic approaches to minimizing its negative effect on search results. Harnessing the expertise of both client-users and professional searchers, we establish principled methods to generate a test collection for identifying and handling inexact duplicate documents. @inproceedings{1009131, author = {Jack G. Conrad and Cindy P. Schriber}, title = {Constructing a text corpus for inexact duplicate detection}, booktitle = {SIGIR '04: Proceedings of the 27th annual international conference on Research and development in information retrieval}, year = {2004}, isbn = {1-58113-881-4}, pages = {582--583}, location = {Sheffield, United Kingdom}, doi = {http://doi.acm.org/10.1145/1008992.1009131}, publisher = {ACM Press}, } ![]() ©2005 Association for Computing Machinery |