![]() ![]() ![]() |
![]() |
|
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Return to Session IXa: Machine learning for IR Topic tracking is complicated when the stories in the stream occur in multiple languages. Typically, researchers have trained only English topic models because the training stories have been provided in English. In tracking, non-English test stories are then machine translated into English to compare them with the topic models. We propose a native language hypothesis stating that comparisons would be more effective in the original language of the story. We first test and support the hypothesis for story link detection. For topic tracking the hypothesis implies that it should be preferable to build separate language-specific topic models for each language in the stream. We compare different methods of incrementally building such native language topic models. @inproceedings{1009061, author = {Leah S. Larkey and Fangfang Feng and Margaret Connell and Victor Lavrenko}, title = {Language-specific models in multilingual topic tracking}, booktitle = {SIGIR '04: Proceedings of the 27th annual international conference on Research and development in information retrieval}, year = {2004}, isbn = {1-58113-881-4}, pages = {402--409}, location = {Sheffield, United Kingdom}, doi = {http://doi.acm.org/10.1145/1008992.1009061}, publisher = {ACM Press}, } ![]() ©2005 Association for Computing Machinery |