Accelerating Substructure Similarity Search for Formula Retrieval

Accelerating Substructure Similarity Search for Formula Retrieval
复制标题

DOI:
10.1007/978-3-030-45439-5_47
复制
发表时间:
2020-03-17
期刊:
Advances in Information Retrieval
影响因子:
--
通讯作者:
Zanibbi R
Zanibbi R
中科院分区:
其他
文献类型:
--
作者:
Zhong W;Rohatgi S;Wu J;Giles CL;Zanibbi R

文献摘要

被引文献

相似文献

使用子结构匹配的公式检索系统是有效的,但遭受由结构匹配的复杂性引起的缓慢的检索时间。我们提出了一个专门的倒排索引和排名安全的动态修剪算法,更快的子结构检索。公式从其运算符树(OPT)表示中索引。我们的模型使用NTCIR-12维基百科公式浏览任务和一个新的公式语料库从数学StackExchange帖子进行评估。我们的方法保留了结构匹配的有效性,同时允许实时执行查询。
Formula retrieval systems using substructure matching are effective, but suffer from slow retrieval times caused by the complexity of structure matching. We present a specialized inverted index and rank-safe dynamic pruning algorithm for faster substructure retrieval. Formulas are indexed from their Operator Tree (OPT) representations. Our model is evaluated using the NTCIR-12 Wikipedia Formula Browsing Task and a new formula corpus produced from Math StackExchange posts. Our approach preserves the effectiveness of structure matching while allowing queries to be executed in real-time.