Efficient Algorithm for Math Formula Semantic Search

Efficient Algorithm for Math Formula Semantic Search
复制标题

DOI:
10.1587/transinf.2015dap0023
复制
发表时间:
2016-04
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Shunsuke Ohashi;Giovanni Yoko Kristianto;Goran Topic;Akiko Aizawa
Shunsuke Ohashi;Giovanni Yoko Kristianto;Goran Topic;Akiko Aizawa
中科院分区:
其他
文献类型:
--
作者:
Shunsuke Ohashi;Giovanni Yoko Kristianto;Goran Topic;Akiko Aizawa

文献摘要

相似文献

数学公式在许多科学领域发挥着重要作用。不管数学公式搜索的重要性如何,传统的基于关键字的检索方法不足以搜索结构为树的数学公式。科学文章中数学公式的数量不断增加以及结构复杂性导致需要大规模的结构感知公式搜索技术。在本文中,我们制定了三种类型的度量来代表数学公式语义相似性的独特特征,并开发了有效的基于哈希的算法来进行近似计算。我们使用 NTCIR-11 Math-2 Task 数据集(一个包含约 6000 万个公式的大规模数学信息检索测试集)进行的实验表明,所提出的方法提高了搜索精度,同时保持了较高的可扩展性和运行效率。关键词: 树哈希, MathML, 数学公式搜索, 信息检索
Mathematical formulae play an important role in many scientific domains. Regardless of the importance of mathematical formula search, conventional keyword-based retrieval methods are not sufficient for searching mathematical formulae, which are structured as trees. The increasing number as well as the structural complexity of mathematical formulae in scientific articles lead to the necessity for large-scale structureaware formula search techniques. In this paper, we formulate three types of measures that represent distinctive features of semantic similarity of math formulae, and develop efficient hash-based algorithms for the approximate calculation. Our experiments using NTCIR-11 Math-2 Task dataset, a large-scale test collection for math information retrieval with about 60million formulae, show that the proposed method improves the search precision while also keeps the scalability and runtime efficiency high. key words: tree hashing, MathML, mathematical formula search, information retrieval