Planning and Factorization for Graph Database Query Optimization and Evaluation
Planning and Factorization for Graph Database Query Optimization and Evaluation
批准号:
RGPIN-2022-04548
负责人:
Godfrey, Parke
金额:
$1.75万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
只要人们看到星星,文化就会把它们联系起来。一些星群如此紧密相连,如此引人注目,人们赋予这些星座意义,将它们与他们世界中的事物联系起来。这并不奇怪。这种模拟世界的方式相当自然。我们描绘出世界上的事物是如何相互关联的,它们之间的“概念线”。我们在脑海中走着这些我们构建的概念图来回答关于我们世界的问题,并进行概括,从我们的地图中提取知识。这也是一个有用的模型,用于为计算机捕获信息,我们称之为数据库。事物,实体,是我们想要存储信息的任何东西。我们可以在相关的实体(事物)之间画边,标记它们的关联方式。我们的数据库是一个很大的图。然后用正确的语言提出问题、查询,我们就可以在需要的时候从图中提取有用的信息。即计算机遍历图以查找所查询的信息。假设我们为接触追踪构建了一个图形数据库。人是实体。当两个人有亲密的身体接触时,我们在图上画了一条标记的边。注意,两个人实体之间可能有很多边,每次他们靠近时都有一条边。每条边都可以用它们的时间来标记。这样的图表数据库可以用于接触者追踪:每当有人生病时,我们可以查询图表来发现他们在一段时间内与谁有过接触。图形数据库对于以自然的方式收集数据非常有用。我们开始看到一些组织建立了非常大的图形数据库。例如,Uniprot SPARQL端点(数据集)最近由63,376,853,475条边组成。Uniprot(通用蛋白质资源)是一个免费访问的,流行的蛋白质数据存储库,生物学家使用。但是这样的图形数据库只有在我们可以使用它们获取有用信息的情况下才有用。由于图形数据库的巨大规模,对它们的查询可能很难回答,也很难有效地评估。如果你需要等待数年才能得到答案,那么一个查询是没有用的!图形数据库是一项相当新的技术。所以我们才刚刚开始学习如何有效地使用和使用它们。这包括如何有效地评估对它们的查询。其他类型的数据库存在的时间更长;我们对如何与他们合作有更深刻的经验。在这项工作中,我们调整了查询优化方法,以提高图数据库的查询评估效率。图数据库的连接查询(一个有用的类)本身就是小图。我们也将查询的答案建模为一个图,一个答案图。这是一种看起来很有希望用于图数据的因子分解技术。让图形数据库真正有用。
英文摘要
As long as people have looked at the stars, cultures have imaged lines connecting them. Some groups of so connected stars so stood out, people ascribed meaning to these constellations, relating them to things within their world. This is not so surprising. This way of modelling the world comes rather naturally. We map out how things in our world are related, the "conceptual lines" between them. We mentally walk these conceptual graphs we have constructed to answer questions about our world, and to generalize, to extract knowledge from our map. This is a useful model to use to capture information for computers, too, what we call a database. The things, entities, are whatever we are wanting to store information about. And we can put edges between entities (things) that are related, labelled with how they are related. So our database is a large graph. Then with the right language to ask questions, queries, we can extract useful information from the graph when needed. That is the computer walking the graph to find the queried information. Imagine we build a graph database for contact tracing. People are the entities. And we put a labelled edge into the graph between two people whenever the two come into close physical contact. Note two people entities might have lots of edges between them, one for each time they came in close proximity. Each edge can be labelled with the time that they did. Such a graph database could be used for contact tracing: whenever someone became ill, we could query the graph to discover with whom they had been in contact over a range of time. Graph databases are quite useful for collecting data in a natural way. And we are beginning to see organizations build extremely large graph databases. For example, the Uniprot SPARQL Endpoint (dataset) consists of 63,376,853,475 edges as of a recent time. Uniprot (UNIversal PROTein resource) is a freely accessible, popular repository of protein data used by biologists. But such graph databases are only as useful as, well, however we can use them to get useful information back out. Because of the immense size graph databases can be, queries over them can be hard to answer, to evaluate efficiently. A query is not useful if you have to wait years for the answer! Graph databases are a fairly new technology. So we are at the start of learning how to use and work with them efficiently. This includes how to evaluate queries over them efficiently. Other types of databases have been around much longer; we have much deeper experience how to work with them. In this work, we adapt methods of query optimization for more efficient query evaluation for graph databases. Conjunctive queries (a useful class) for graph databases are small graphs themselves. We model the answers to a query as a graph too, an answer graph. This is a _factorization_ technique that looks to be quite promising for graph data. And so towards making graph databases truly useful.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Cost-based Optimization for SPARQL Property Path Queries
-
批准号:RGPIN-2015-04242
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2019
-
负责人:Godfrey, Parke
-
依托单位:
Cost-based Optimization for SPARQL Property Path Queries
-
批准号:RGPIN-2015-04242
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2018
-
负责人:Godfrey, Parke
-
依托单位:
Building data visualization and exploration support into an embedded database system: Big Data in the small
-
批准号:461932-2013
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$3.82万
-
财政年份:2017
-
负责人:Godfrey, Parke
-
依托单位:
Cost-based Optimization for SPARQL Property Path Queries
-
批准号:RGPIN-2015-04242
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2017
-
负责人:Godfrey, Parke
-
依托单位:
Cost-based Optimization for SPARQL Property Path Queries
-
批准号:RGPIN-2015-04242
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2016
-
负责人:Godfrey, Parke
-
依托单位:
Building data visualization and exploration support into an embedded database system: Big Data in the small
-
批准号:461932-2013
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$3.82万
-
财政年份:2016
-
负责人:Godfrey, Parke
-
依托单位:
Cost-based Optimization for SPARQL Property Path Queries
-
批准号:RGPIN-2015-04242
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2015
-
负责人:Godfrey, Parke
-
依托单位:
Database support for 4D virtual and mirror spaces
-
批准号:228108-2009
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2014
-
负责人:Godfrey, Parke
-
依托单位:
Database support for 4D virtual and mirror spaces
-
批准号:228108-2009
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2012
-
负责人:Godfrey, Parke
-
依托单位:
Database support for 4D virtual and mirror spaces
-
批准号:228108-2009
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2011
-
负责人:Godfrey, Parke
-
依托单位:
Database support for 4D virtual and mirror spaces
-
批准号:228108-2009
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2010
-
负责人:Godfrey, Parke
-
依托单位:
Database support for 4D virtual and mirror spaces
-
批准号:228108-2009
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2009
-
负责人:Godfrey, Parke
-
依托单位:
Relational support for preference queries and cooperative answers
-
批准号:228108-2004
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.8万
-
财政年份:2008
-
负责人:Godfrey, Parke
-
依托单位:
Relational support for preference queries and cooperative answers
-
批准号:228108-2004
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.8万
-
财政年份:2007
-
负责人:Godfrey, Parke
-
依托单位:
Relational support for preference queries and cooperative answers
-
批准号:228108-2004
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.8万
-
财政年份:2006
-
负责人:Godfrey, Parke
-
依托单位:
Relational support for preference queries and cooperative answers
-
批准号:228108-2004
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.8万
-
财政年份:2005
-
负责人:Godfrey, Parke
-
依托单位:
Relational support for preference queries and cooperative answers
-
批准号:228108-2004
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.8万
-
财政年份:2004
-
负责人:Godfrey, Parke
-
依托单位:
Tools and methodologies for reasoning about database knowledge
-
批准号:228108-2000
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.24万
-
财政年份:2003
-
负责人:Godfrey, Parke
-
依托单位:
Tools and methodologies for reasoning about database knowledge
-
批准号:228108-2000
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.24万
-
财政年份:2002
-
负责人:Godfrey, Parke
-
依托单位:
Tools and methodologies for reasoning about database knowledge
-
批准号:228108-2000
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.24万
-
财政年份:2001
-
负责人:Godfrey, Parke
-
依托单位:
海外基金