Welcome to D
SIGMOD 2004
PODS 2004
SIGMOD RECOR
CIKM 2004
DASFAA 2004
DBPL 2003
DE-BULLETIN
DEBS 2004
DMKD 2004
DMSN 2004
DOLAP 2004
DPDJ 2004
EDBT 2004
ER 2003
GIS 2004
HDP 2004
HYPERTEXT 20
ICDE 2004
ICDT 2003
JCDL 2004
MDM
MIR 2004
MIS 2004
MMDB 2004
MOBIDE 2003
RIDE 2004
SBBD 2003
SIGIR FORUM
SIGIR 2004
<<< = SIGIR'04 Pap>>>
SIGKDD EXPLO
SIGKDD 2004
SSDBM 2004
SSTD 2003
TIME 2004
TODS 2004
VLDB 2004
VLDB Journal
WEBDB 2004
WIDM 2004
XIME-P 2004
Footer

Corpus Structure, Language Models, and Ad Hoc Information Retrieval


Oren Kurland and Lillian Lee

  View Paper (PDF)  

Return to Session Va: Language models


Abstract

Most previous work on the recently developed language-modeling approach to information retrieval focuses on document-specific characteristics, and therefore does not take into account the structure of the surrounding corpus. We propose a novel algorithmic framework in which information provided by document-based language models is enhanced by the incorporation of information drawn from clusters of similar documents. Using this framework, we develop a suite of new algorithms. Even the simplest typically outperforms the standard language-modeling approach in precision and recall, and our new interpolation algorithm posts statistically significant improvements for both metrics over all three corpora tested.

BIBTEX


@inproceedings{1009027,   author = {Oren Kurland and Lillian Lee},
  title = {Corpus structure, language models, and ad hoc information retrieval},
  booktitle = {SIGIR '04: Proceedings of the 27th annual international conference on Research and development in information retrieval},
  year = {2004},
  isbn = {1-58113-881-4},
  pages = {194--201},
  location = {Sheffield, United Kingdom},
  doi = {http://doi.acm.org/10.1145/1008992.1009027},
  publisher = {ACM Press},
  
}



©2005 Association for Computing Machinery