Semantic Index : Scalable Query Answering without Forward Chaining or Exponential Rewritings

Semantic Index : Scalable Query Answering without Forward Chaining or Exponential Rewritings
复制标题

语义索引:无需前向链接或指数重写的可扩展查询应答

DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
Diego Calvanese
Diego Calvanese
中科院分区:
--
文献类型:
--
作者:
M. Rodriguez;Diego Calvanese

文献摘要

被引文献

相似文献

需要制度为SPARQL 1.1中的丰富推论提供了支持。这极大地有助于在应用中使用推理。在这种情况下,语义网(SW)社区的一个特殊兴趣是基于本体的数据访问(OBDA),即,通过本体论的词汇和语义来查询大量主张数据。在过去的几年中,OBDA引起了很多关注,尽管在理论方面取得了进步,例如,OWL 2 QL的定义,实现了大型本体和大数据集的有效且可扩展的推理仍然是有问题的。在这种情况下,查询答案的最广泛的推理技术是使用前向链接的推论实现。该技术具有多个优点,例如,所有推论都是离线完成的,它相对容易实现,并且在查询时间提供了高性能。但是,如果本体论的术语部分很大,则物质化可能需要很长时间,并且可能会大大增加应用程序的存储要求。在几种相关用例中,这些缺点可能会使这一技术不受欢迎或不可行。实现化的替代方法是查询通过查询重写,其中所有推理都在线完成。查询重写通常被提升为查询大量数据的最有效方法。但是,实际上,我们尚未看到这些技术的广泛采用。主要原因是通过重写产生的查询通常太大或太复杂。例如,在大猫头鹰2 QL本体论的情况下,重写通常会产生数百或数千个子征服。在本文中,我们提出了一种将离线和在线推理结合在一起的技术,以避免上述实体化和查询重写和保证,在实践中保证了最小的三重商店的时间和空间以及快速查询的响应。我们使用RDBMS系统作为数据后端制定技术;但是,我们注意到该技术很容易适应天然三重储物。同样,在以下内容中,我们专注于OWL 2 QL直接支出制度,但是,该技术也可以与RDFS制度一起使用。语义索引。语义索引技术的核心思想是将OWL 2 QL本体术语的所需层次结构(即Tbox t)编码为我们分配给类和属性的数字索引。我们使用这些值来插入本体论的主张数据,即abox a,并使用范围查询来检索层次结构和Abox主张所需要的三元组。这使我们能够创建三重存储库,这些存储库几乎是原始数据的大小,并且已经编码了本体学的大多数语义。结合一种简单的重写技术,我们能够为SPARQL 1.1 ABOX查询提供快速,可扩展的查询答案,而OWL 2 QL QL构成态度,同时保持健全性和完整性。我们的建议与管理大型及时关系的技术密切相关
Entailment regimes add support for rich inferences in SPARQL 1.1. This greatly facilitates the use of reasoning in applications. In this context, one special interest of the Semantic Web (SW) community is Ontology Based Data Access (OBDA), i.e., querying large volumes of assertional data through the vocabulary and semantics of ontologies. OBDA has recieved a lot of attention in the last years, however, while there have been advances on the theoretical side, e.g., the definition of OWL 2 QL, realizing efficient and scalable reasoning for large ontologies and large data sets is still problematic. In this context, the most widespread reasoning technique for query answering is the materialization of inferences using forward chaining. This technique has several advantages, e.g., all inferences are done off-line, it is relatively easy to implement, and it offers high-performance at query time. However, if the terminological part of the ontology is large, materialization may require a long time and may considerably increase the storage requirements of the application. These disadvantages can turn this technique undesirable or unfeasible in several relevant use cases. An alternative to materialization is query answering by query rewriting, in which all reasoning is done on-line. Query rewriting has often been promoted as the most efficient way to query large volumes of data. However, in practice we have not seen a widespread adoption of these techniques. The main reason is that the queries generated by rewriting are often too large or too complex; for example, in the case of large OWL 2 QL ontologies, rewritings often generate hundreds or thousands of subqueries. In this paper we present a technique that combines off-line and on-line reasoning to avoid the aforementioned issues of materialization and query rewriting and guaranteeing, in practice, minimal time and space for the construction of the triple store and fast query answering. We formulate the technique using RDBMS systems as the data backend; however, we note that the technique can easily be adapted to native triple stores. Likewise, in the following we focus on the OWL 2 QL direct entailment regime, however, the technique can also be used with the RDFS regime. Semantic Index. The core idea of the semantic index technique is to encode the entailed hierarchies of the terminology of the OWL 2 QL ontology, i.e., the TBox T , into numeric indexes that we assign to classes and properties. We use these values to insert the assertional data of the ontology, i.e., the ABox A, into the DB, and use range queries to retrieve the triples entailed by the hierarchies and the ABox assertions. This allows us to create triple repositories that are almost the size of the original data and already encode most of the semantics of the ontology. Combined with a simple rewriting technique, we are able to provide fast and scalable query answering for SPARQL 1.1 ABox queries under the OWL 2 QL entailment regime while preserving soundness and completeness. Our proposal is strongly related to techniques for managing large transitive relations in