Welcome to D
SIGMOD 2004
PODS 2004
SIGMOD RECOR
CIKM 2004
DASFAA 2004
DBPL 2003
DE-BULLETIN
DEBS 2004
DMKD 2004
DMSN 2004
DOLAP 2004
DPDJ 2004
EDBT 2004
ER 2003
GIS 2004
HDP 2004
HYPERTEXT 20
ICDE 2004
ICDT 2003
JCDL 2004
MDM
MIR 2004
MIS 2004
MMDB 2004
MOBIDE 2003
RIDE 2004
SBBD 2003
SIGIR FORUM
SIGIR 2004
<<< = SIGIR'04 Pap>>>
SIGKDD EXPLO
SIGKDD 2004
SSDBM 2004
SSTD 2003
TIME 2004
TODS 2004
VLDB 2004
VLDB Journal
WEBDB 2004
WIDM 2004
XIME-P 2004
Footer

Query-related data extraction of hidden web documents


Yih-Ling Hedley, Muhammad Younas, A. James, and M. Sanderson

  View Paper (PDF)  

Return to Posters


Abstract

The larger amount of information on the Web is stored in document databases and is not indexed by general-purpose search engines (i.e., Google and Yahoo). Such information is dynamically generated through querying databases - which are referred to as Hidden Web databases. Documents returned in response to a user query are typically presented using template-generated Web pages. This paper proposes a novel approach that identifies Web page templates by analysing the textual contents and the adjacent tag structures of a document in order to extract query-related data. Preliminary results demonstrate that our approach effectively detects templates and retrieves data with high recall and precision.

BIBTEX


@inproceedings{1009119,   author = {Y. L. Hedley and M. Younas and A. James and M. Sanderson},
  title = {Query-related data extraction of hidden web documents},
  booktitle = {SIGIR '04: Proceedings of the 27th annual international conference on Research and development in information retrieval},
  year = {2004},
  isbn = {1-58113-881-4},
  pages = {558--559},
  location = {Sheffield, United Kingdom},
  doi = {http://doi.acm.org/10.1145/1008992.1009119},
  publisher = {ACM Press},
  
}



©2005 Association for Computing Machinery