Batch Processing of Top-k Spatial-Textual Queries

Batch Processing of Top-k Spatial-Textual Queries
复制标题

Top-k 空间文本查询的批处理

DOI:
10.1145/3196155
复制
发表时间:
2018-05
影响因子:
1.9
通讯作者:
Timos Sellis
Timos Sellis
中科院分区:
--
文献类型:
--
作者:
Farhana M Choudhury;J Shane Culpepper;Zhifeng Bao;Timos Sellis

文献摘要

参考文献

被引文献

相似文献

自2000年代中期以来,已经提出了多种索引技术来有效地回答top-k空间文本查询。然而,所有这些方法都集中在一次回答一个查询。相比之下,如何设计有效的算法,可以利用传入查询之间的相似性,以提高性能却很少受到关注。在这篇文章中,我们研究了一系列有效的方法来批处理多个top-k空间文本查询并发。我们仔细设计了各种索引结构的问题空间,探索优先空间和文本属性对系统性能的影响。具体来说,我们提出了一个有效的遍历方法,SF-Sep,在现有的空间优先级的索引结构。然后,我们提出了一个新的空间优先级的索引结构,MIR-Tree支持过滤和细化为基础的技术,SF-Grp。为了支持文本密集型数据的处理,我们提出了一个增强的,倒排索引结构,可以很容易地添加到现有的文本搜索引擎架构和一种新的遍历方法的批处理的查询。在所有这些方法中,目标都是通过分担类似查询的I/O成本来提高整体性能。最后,我们证明了显着的I/O节省在我们的算法比传统的方法通过广泛的实验在三个真实的数据集,并比较不同的数据集的属性如何影响性能。流媒体、连续查询的微处理器和隐私感知搜索中的许多应用都可以从这一系列工作中受益。
Since the mid-2000s, everal indexing techniques have been proposed to efficiently answer top-k spatial-textual queries. However, all of these approaches focus on answering one query at a time. In contrast, how to design efficient algorithms that can exploit similarities between incoming queries to improve performance has received little attention. In this article, we study a series of efficient approaches to batch process multiple top-k spatial-textual queries concurrently. We carefully design a variety of indexing structures for the problem space by exploring the effect of prioritizing spatial and textual properties on system performance. Specifically, we present an efficient traversal method, SF-Sep, over an existing space-prioritized index structure. Then, we propose a new space-prioritized index structure, the MIR-Tree to support a filter-and-refine based technique, SF-Grp. To support the processing of text-intensive data, we propose an augmented, inverted indexing structure that can easily be added into existing text search engine architectures and a novel traversal method for batch processing of the queries. In all of these approaches, the goal is to improve the overall performance by sharing the I/O costs of similar queries. Finally, we demonstrate significant I/O savings in our algorithms over traditional approaches by extensive experiments on three real datasets and compare how properties of different datasets affect the performance. Many applications in streaming, micro-batching of continuous queries, and privacy-aware search can benefit from this line of work.
DOI: 10.1109/perser.2005.1506394
发表时间: 2005-07
期刊: ICPS '05. Proceedings. International Conference on Pervasive Services, 2005.
影响因子: --
作者:
H. Kido;Y. Yanagisawa;T. Satoh
通讯作者: H. Kido;Y. Yanagisawa;T. Satoh
DOI: 10.1145/2786006.2786008
发表时间: 2015-05
期刊: Second International ACM Workshop on Managing and Mining Enriched Geo-Spatial Data
影响因子: --
作者:
F. Choudhury;J. Culpepper;T. Sellis
通讯作者: F. Choudhury;J. Culpepper;T. Sellis
DOI: 10.1145/956863.956944
发表时间: 2003-11
期刊: --
影响因子: --
作者:
A. Broder;David Carmel;Michael Herscovici;A. Soffer;Jason Y. Zien
通讯作者: A. Broder;David Carmel;Michael Herscovici;A. Soffer;Jason Y. Zien
DOI: 10.1145/2433396.2433412
发表时间: 2013-02
期刊: Proceedings of the sixth ACM international conference on Web search and data mining
影响因子: --
作者:
C. Dimopoulos;Sergey Nepomnyachiy;Torsten Suel
通讯作者: C. Dimopoulos;Sergey Nepomnyachiy;Torsten Suel
DOI: 10.1016/0306-4573(95)00020-h
发表时间: 1995-11
期刊: Inf. Process. Manag.
影响因子: --
作者:
Howard R. Turtle;James Flood
通讯作者: Howard R. Turtle;James Flood