Monolingual and Crosslingual SMS-based FAQ Retrieval

Monolingual and Crosslingual SMS-based FAQ Retrieval
复制标题

基于短信的单语和跨语言常见问题解答检索

DOI:
10.1145/2701336.2701634
复制
发表时间:
2013
期刊:
Proceedings of the 4th and 5th Annual Meetings of the Forum for Information Retrieval Evaluation
影响因子:
--
通讯作者:
Johannes Leveling
Johannes Leveling
中科院分区:
--
文献类型:
--
作者:
Johannes Leveling

文献摘要

被引文献

相似文献

本文介绍了DCU第二次参与FIRE基于短信的常见问题检索任务的结果。对于FIRE 2012,我们提交了单语言英语和印地语以及跨语言英语到印地语子任务的运行。与FIRE 2011的实验相比,我们的系统通过使用单个检索引擎(而不是三个)和使用单一方法来检测域外查询(而不是三个)而得到了简化。在我们的方法中,SMS查询被转换为规范化的更正表单,并提交给检索引擎,以获得FAQ结果的排名列表。分类器根据从训练数据中提取的特征进行训练,然后确定哪些查询超出了域,哪些没有。对于我们的英语到印地语的跨语言实验,我们训练了一个统计机器翻译系统,用于印地语到英语的翻译,将完整的印地语FAQ文档翻译成英语。然后,检索操作更正后的英语输入,并从翻译后的印地语FAQ文档中检索结果。我们最好的实验实现了单语英语子任务的MRR为0.949,单语印地语子任务的MRR为0.880,跨语子任务的MRR为0.450。
This paper presents results for DCU's second participation in the SMS-based FAQ Retrieval task at FIRE. For FIRE 2012, we submitted runs for the monolingual English and Hindi and the crosslingual English to Hindi subtasks. Compared to our experiments for FIRE 2011, our system was simplified by using a single retrieval engine (instead of three) and using a single approach for detection of out of domain queries (instead of three). In our approach, the SMS queries are transformed into a normalized, corrected form and submitted to a retrieval engine to obtain a ranked list of FAQ results. A classifier trained on features extracted from the training data then determines which queries are out of domain and which are not. For our crosslingual English to Hindi experiments, we trained a statistical machine translation system for Hindi to English translation to translate the full Hindi FAQ documents into English. The retrieval then operates on the corrected English input and retrieves results from the translated Hindi FAQ documents. Our best experiments achieved an MRR of 0.949 for the monolingual English subtask, 0.880 for the monolingual Hindi subtask, and 0.450 for the crosslingual subtask.