Large Language Models for Query Optimisation: A New Paradigm in Database Systems
Large Language Models for Query Optimisation: A New Paradigm in Database Systems
批准号:
2726025
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Research ImpactThe vision of this research project is to revolutionize query optimisation in database (DB) systems using Large Language Models (LLMs). LLMs belong to the class of foundation models, AI paradigms capable of tackling multiple downstream tasks. I propose a comprehensive investigation into the ability of LLMs to act as 'brains' for efficient query processing in DB systems. Constructing queries efficiently is essential for DBs to run quickly. Database systems leverage query rewriting algorithms to transform queries, so they execute with low latency. Traditionally, this is done through the application of manually constructed rewrite rules, where the rewritten query should yield equivalent output as the original one, while exhibiting higher performance. Replacing white-box query optimisation strategies with zero-shot and few-shot learning via LLMs is an important step towards autonomous DB systems. Success in this research will be impactful in the database community, as it will prove that automating the end-to-end query rewrite process by bridging knowledge from natural language processing and database systems is viable. As an example, submitting a relational query to the DB will require no ad-hoc optimisation from the database administrator (DBA), in the presence of LLM-generated rewrite rules. Furthermore, I expect vendors to save hundreds of developer hours spent on extending the existing systems with ever-more rewrite rules. The necessity for new query rewrite rules is driven by changes in the queries executed, including non-human transactions such as those generated by web applications.Aims and ObjectivesThe envisioned goals of this research are to investigate and prove the following:1. The ability of LLMs to 'understand' the intricacies of existing DB systems. By capturing the logical and physical facets of current DBs, I envision LLMs to adapt well to various downstream database optimisation tasks.2. The efficiency of LLMs as query rewrite mechanisms. This goal aims to uncover how fast (i.e., zero-shot, few-shot) LLMs can learn to optimise queries and their performance against existing DBs. 3. The assets required to build an LLM-powered DB system. This objective aims to reduce the complexity of integrating LLMs into DB systems and beyond to a range of software systems and algorithms. MethodologyThe initial research methodology is to establish an LLM-based foundation for automating query rewriting. The deliverables will serve as artifacts to tackle tasks beyond query optimisation. There are two constituent parts. First, a pipeline for guiding the application of rewrite rules for DB queries will be implemented. The purpose of this is to ensure the order in which the rules are applied is optimal. Generally, finding the optimal order of applying query rewrite rules is an NP-hard problem. The reason is that applying a suboptimal rewrite rule early in the chain may prevent globally optimal rule applications. Second, the rewrite rules are generally designed by human experts, so instead, generating query rewrite rules via LLMs will be investigated through the prism of prompt engineering and adapters to eliminate human error and guesswork. EPSRC Strategic AlignmentBridging LLMs and DBs brings the research community closer to an autonomous DB and it's aligned with the "Artificial intelligence (AI), digitalisation and data: driving value and security" EPSRC objective. Serving information through natural language processing presents a real opportunity for driving innovation in the UK technology sector, with an important economic impact.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金