Efficient hashing with lookups in two memory accesses

Efficient hashing with lookups in two memory accesses
复制标题

在两次内存访问中进行查找的高效散列

DOI:
--
复制
发表时间:
2004
期刊:
ACM-SIAM Symposium on Discrete Algorithms
影响因子:
--
通讯作者:
R. Panigrahy
R. Panigrahy
中科院分区:
--
文献类型:
--
作者:
R. Panigrahy

文献摘要

被引文献

相似文献

哈希的研究与Azar等人的分析密切相关。 ,此大幅度降低了垃圾箱的最大负载。哈希,由于一个物品可以放在两个水桶之一中,因此我们可以使用此事实来减少最大负载后移动一个项目。 2,具有很高的概率,同时可以实现高内存利用率,即使两个项目的空间是预先分配的,也可以在硬件实现中所需可以存储高内存利用率的项目。 log n)时间及以上log log n+o(1)移动,概率很高,并且在期望中持续不断的时间。•保持83.75%的内存利用率,而无需在插入过程中进行动态分配。分析插入过程中执行的移动数量与最大载荷的最大载荷之间的权衡。 )。
The study of hashing is closely related to the analysis of balls and bins. Azar et. al. [1] showed that instead of using a single hash function if we randomly hash a ball into two bins and place it in the smaller of the two, then this dramatically lowers the maximum load on bins. This leads to the concept of two-way hashing where the largest bucket contains O(log log n) balls with high probability. The hash look up will now search in both the buckets an item hashes to. Since an item may be placed in one of two buckets, we could potentially move an item after it has been initially placed to reduce maximum load. Using this fact, we present a simple, practical hashing scheme that maintains a maximum load of 2, with high probability, while achieving high memory utilization. In fact, with n buckets, even if the space for two items are pre-allocated per bucket, as may be desirable in hardware implementations, more than n items can be stored giving a high memory utilization. Assuming truly random hash functions, we prove the following properties for our hashing scheme.• Each lookup takes two random memory accesses, and reads at most two items per access.• Each insert takes O(log n) time and up to log log n+O(1) moves, with high probability, and constant time in expectation.• Maintains 83.75% memory utilization, without requiring dynamic allocation during inserts.We also analyze the trade-off between the number of moves performed during inserts and the maximum load on a bucket. By performing at most h moves, we can maintain a maximum load of O(log log n/h log (log log n/h)). So, even by performing one move, we achieve a better bound than by performing no moves at all.