![]() ![]() ![]() |
![]() |
|
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Return to Posters This paper presents a new two-phase pattern (2PP) discovery technique for information extraction. 2PP consists of orthographic pattern discovery (OPD) and semantic pattern discovery (SPD) where the OPD determines the structural features from an identified region of a document and the SPD discovers a dominant semantic pattern for the region via inference, apposition and analogy. Then the discovered pattern is applied back into the region to extract required data items through pattern matching. We evaluated 2PP using 6500 data items and obtained effective result. @inproceedings{1009107, author = {Liping Ma and John Shepherd}, title = {Information extraction using two-phase pattern discovery}, booktitle = {SIGIR '04: Proceedings of the 27th annual international conference on Research and development in information retrieval}, year = {2004}, isbn = {1-58113-881-4}, pages = {534--535}, location = {Sheffield, United Kingdom}, doi = {http://doi.acm.org/10.1145/1008992.1009107}, publisher = {ACM Press}, } ![]() ©2005 Association for Computing Machinery |