An empirical study of tokenization strategies for biomedical information retrieval

An empirical study of tokenization strategies for biomedical information retrieval
复制标题

DOI:
10.1007/s10791-007-9027-7
复制
发表时间:
2007-10-01
期刊:
INFORMATION RETRIEVAL
影响因子:
--
通讯作者:
Zhai, ChengXiang
Zhai, ChengXiang
中科院分区:
其他
文献类型:
--
作者:
Jiang, Jing;Zhai, ChengXiang

文献摘要

被引文献

相似文献

由于生物医学文本中生物名称的多样性,适当的标记化是生物医学信息检索的重要预处理步骤。尽管它的重要性,一直很少有研究对生物医学文本的各种标记化策略的评估。在这项工作中,我们进行了一个仔细的,系统的评估一组标记化的专门文件检索的所有可用的TREC生物医学文本集合,使用两个代表性的检索方法和伪相关反馈方法。我们还研究了词干提取和停用词去除对检索性能的影响。正如预期的那样,我们的实验结果表明,标记化可以显着影响检索精度;适当的标记化可以提高性能高达96%,平均平均精度(MAP)。特别是,它表明,不同的查询类型需要不同的标记化语法,词干是有效的,只有在某些查询,和停止词删除一般不会提高检索性能的生物医学文本。
Due to the great variation of biological names in biomedical text, appropriate tokenization is an important preprocessing step for biomedical information retrieval. Despite its importance, there has been little study on the evaluation of various tokenization strategies for biomedical text. In this work, we conducted a careful, systematic evaluation of a set of tokenization heuristics on all the available TREC biomedical text collections for ad hoc document retrieval, using two representative retrieval methods and a pseudo-relevance feedback method. We also studied the effect of stemming and stop word removal on the retrieval performance. As expected, our experiment results show that tokenization can significantly affect the retrieval accuracy; appropriate tokenization can improve the performance by up to 96%, measured by mean average precision (MAP). In particular, it is shown that different query types require different tokenization heuristics, stemming is effective only for certain queries, and stop word removal in general does not improve the retrieval performance on biomedical text.