An enhanced dynamic hash TRIE algorithm for lexicon search

An enhanced dynamic hash TRIE algorithm for lexicon search
复制标题

DOI:
10.1080/17517575.2012.665483
复制
发表时间:
2012-11
影响因子:
4.4
通讯作者:
Lai Yang;Lida Xu;Zhongzhi Shi
Lai Yang;Lida Xu;Zhongzhi Shi
中科院分区:
计算机科学3区
文献类型:
--
作者:
Lai Yang;Lida Xu;Zhongzhi Shi

文献摘要

被引文献

相似文献

沿着订单、客户和材料的不断增长,信息检索(IR)对企业系统至关重要。本文提出了一种增强型动态哈希TRIE(eDH-TRIE)算法,可用于中日韩(CJK)分词词典检索和URL识别。特别地,eDH-TRIE算法适用于Unicode检索。提出了Auto-Array算法和Hash-Array算法来处理辅助内存分配,前者根据需要改变大小,无需冗余重组,后者用数组代替链表,节省了内存开销。对比实验表明,Auto-Array算法和Hash-Array算法具有更好的空间性能,可以应用于多种场合。eDH-TRIE的速度和存储进行了评估,并与天真的DH-TRIE算法进行了比较。实验结果表明,eDH-TRIE算法具有较好的性能。这些算法减少了内存开销,提高了IR速度。
Information retrieval (IR) is essential to enterprise systems along with growing orders, customers and materials. In this article, an enhanced dynamic hash TRIE (eDH-TRIE) algorithm is proposed that can be used in a lexicon search in Chinese, Japanese and Korean (CJK) segmentation and in URL identification. In particular, the eDH-TRIE algorithm is suitable for Unicode retrieval. The Auto-Array algorithm and Hash-Array algorithm are proposed to handle the auxiliary memory allocation; the former changes its size on demand without redundant restructuring, and the latter replaces linked lists with arrays, saving the overhead of memory. Comparative experiments show that the Auto-Array algorithm and Hash-Array algorithm have better spatial performance; they can be used in a multitude of situations. The eDH-TRIE is evaluated for both speed and storage and compared with the naïve DH-TRIE algorithms. The experiments show that the eDH-TRIE algorithm performs better. These algorithms reduce memory overheads and speed up IR.