SMS based FAQ Retrieval for Hindi, English and Malayalam

SMS based FAQ Retrieval for Hindi, English and Malayalam
复制标题

基于短信的印地语、英语和马拉雅拉姆语常见问题解答检索

DOI:
10.1145/2701336.2701642
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
Rahis Shaikh
Rahis Shaikh
中科院分区:
农林科学3区
文献类型:
--
作者:
A. Shaikh;R. Shah;Rahis Shaikh

文献摘要

被引文献

相似文献

本文介绍了我们在 FIRE 2012 和 FIRE 2013 中基于短信的常见问题检索单语任务的方法。与我们之前在 FIRE 2011 中提交的该任务的解决方案相比,当前的方法可以更准确地预测短信和常见问题的匹配。这次除了印地语和英语之外,我们还提供马拉雅拉姆语(印度语言)的短信和常见问题匹配解决方案。为了执行短信查询和常见问题解答数据库之间的匹配,我们引入了增强的相似度得分、接近度得分、增强的长度得分和答案匹配系统。我们引入术语词干提取,并考虑在短信查询和常见问题解答中加入相邻术语的效果,以提高相似度得分。我们提出了一种新颖的方法来标准化常见问题解答和短信令牌,以提高印地语的准确性。此外,我们建议使用一些字符替换来处理 SMS 查询中的错误。我们通过考虑 FIRE 提供的来自健康、电信、保险和铁路预订等多个不同领域的许多现实常见问题解答数据集来证明我们方法的有效性。实验结果证实,我们针对基于短信的常见问题解答检索单语任务的解决方案非常令人鼓舞,并且在英语、印地语和马拉雅拉姆语的提交中表现非常出色。对于 FIRE 2012 中基于英语、印地语和马拉雅拉姆语 SMS 的常见问题检索单语任务,我们的方法的平均倒数排名 (MRR) 分数分别为 0.971、0.973 和 0.761。此外,我们的解决方案在 FIRE 2013 中的 MRR 分数等于 0.971 的印地语任务中名列前茅。我们的方法在 FIRE 中对于英语语言也表现良好。 2013 年,尽管语音查询的记录与普通短信查询一起包含在测试数据集中。
This paper presents our approach for the SMS-based FAQ Retrieval monolingual task in FIRE 2012 and FIRE 2013. Current approach predicts the matching of an SMS and FAQs more accurately as compared to our previous solution for this task which was submitted in FIRE 2011. We provide solution for SMS and FAQs matching in Malayalam language (an Indian language) in addition to Hindi and English this time. In order to perform a matching between SMS queries and FAQ database, we introduce enhanced similarity score, proximity score, enhanced length score and an answer matching system. We introduce the stemming of terms and consider the effects of joining adjacent terms in SMS query and FAQ to improve the similarity score. We propose a novel method to normalize FAQ and SMS tokens to improve the accuracy for Hindi language. Moreover, we suggest a few character substitutions to handle error in the SMS query. We demonstrate the effectiveness of our approach by considering many real-life FAQ-datasets provided by FIRE from a number of different domains such as Health, Telecom, Insurance and Railway booking. Experimental results confirm that our solution for the SMS-based FAQ Retrieval monolingual task is very encouraging and among the top submissions which performed very well for English, Hindi and Malayalam. The Mean Reciprocal Rank (MRR) scores for our approach are 0.971, 0.973 and 0.761 respectively for English, Hindi and Malayalam SMS-based FAQ Retrieval monolingual task in FIRE 2012. Furthermore, our solution topped the task for Hindi language with MRR score equal to 0.971 in FIRE 2013. Our approach performs very well for English language as well in FIRE 2013 despite transcripts of the speech queries are included in test dataset along with the normal SMS queries.