课题基金 / 基金详情

Automated Classical-to-Quantum Data Encoding for Genomics Data

Automated Classical-to-Quantum Data Encoding for Genomics Data
基因组数据的自动经典到量子数据编码
批准号:
10073337
负责人:
金额:
$6.37万
依托单位:
依托单位国家:
英国
项目类别:
Grant for R&D
财政年份:
2023
资助国家:
英国
项目状态:
已结题
起止时间:
2023 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
生物医学数据是开发机器学习模型的基本资源,可以帮助诊断、治疗和预防疾病。然而,由于生物医学数据的敏感性质和伦理考虑,生物医学数据的收集、存储和共享面临着巨大的挑战。医疗数据受到严格的法规和隐私法的约束,这给研究人员访问和共享数据带来了挑战。此外,在单个数据集上训练的机器学习模型往往过于适合新数据,并且可能不能很好地概括新数据,这限制了它们在现实世界中的潜在应用。为了应对这些挑战,我们必须探索新的生物医学机器学习方法,这种方法可以利用大量和多样化的数据集,同时确保数据隐私和安全。一种有前景的新解决方案是将量子计算的能力与联邦学习(FL)的优点相结合,即混合经典-量子联邦学习。这种分布式量子学习方法使组织能够在各自的经典数据集上训练混合量子机器学习模型,而无需共享原始数据。在联合学习中,机器学习模型训练被分发到可以访问量子处理单元(QPU)的单个设备或服务器来运行量子部分,然后在它们各自的数据集上训练模型。遗憾的是,经典数据集不能直接加载到量子计算机中进行处理,它们需要预先编码成量子计算机可以理解的形式。本质上,经典到量子数据编码是将经典数据转换为量子态以便在量子算法中进一步使用的过程。然而,由于量子比特数量的限制,由于当前一代的量子硬件又称噪声中尺度量子(NISQ)硬件,现有的通用数据编码方法并不总是适合于此目的,特别是当涉及到DNA序列等数据集时。例如,我们最近在DNA序列上测试了一种流行的编码方案,称为“幅度编码”,发现了以下缺点:它对输入数据高度敏感,因为DNA序列的微小变化可能会导致输出的大变化。我们的想法是开发一个高效且自动化的经典到量子数据编码软件即服务(SaaS)工具包,代号为NZ-SeQTech。它针对混合基因组数据和联合量子学习用例进行了优化。NZ-SeQTech的目标是为从事基因组学工作的研究人员、学生、专业人员和爱好者提供一种易于使用和访问的数据编码工具,用于在联合设置中训练混合经典-量子模型。
英文摘要
Biomedical data is an essential resource for developing machine learning models that can aid in diagnosis, treatment, and prevention of diseases. However, the collection, storage, and sharing of biomedical data presents significant challenges due to their sensitive nature and ethical considerations. Healthcare data is subject to strict regulations and privacy laws, making it challenging for researchers to access and share data. Moreover, machine learning models trained on a single dataset tend to overfit, and may not generalise well to new data, which limits their potential use in real-world applications.To address these challenges, we must explore new approaches to biomedical machine learning that can leverage large and diverse datasets, whilst also ensuring data privacy and security. One promising novel solution is to combine the power of quantum computing, with the benefits of federated learning (FL), namely, hybrid classical-quantum federated learning. This distributed quantum learning approach enables organisations to train hybrid quantum machine learning models on their respective classical datasets, without sharing raw data. In federated learning, the machine learning model training is distributed to individual devices or servers with access to quantum processing units (QPUs) to run the quantum part, which then trains the model on their respective datasets.Unfortunately, classical datasets cannot directly be loaded into a quantum computer for processing, they need to be encoded into a form that a quantum computer can understand beforehand. In essence, classical-to-quantum data encoding is the process of converting classical data into quantum states for further usage in a quantum algorithm. However, due to the limitations on the number of qubits, due to the current generation of quantum hardware a.k.a Noisy-Intermediary Scale Quantum (NISQ) hardware, existing general data encoding methods are not always fit for purpose, especially when it comes to datasets such as DNA sequences. For example, we've recently tested a popular encoding scheme known as 'Amplitude encoding' on DNA sequences and found the following shortcomings: it is highly-sensitivity to input data, as small changes in the DNA sequence can result in large changes in the output.Our idea is to develop an efficient and automated classical-to-quantum data encoding software-as-a-service (SaaS) toolkit codenamed NZ-SeQTech. One that is optimised for hybrid genomics data and federated quantum learning use cases. The aim is for NZ-SeQTech to provide; researchers, students, professionals, and enthusiasts working in Genomics with an easy-to-use and accessible data encoding tool for training hybrid classical-quantum models in a federated setup.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金