Data compression for biomedical data analysis
Data compression for biomedical data analysis
批准号:
RGPIN-2022-03074
负责人:
Yu, YunWilliam
金额:
$2.11万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
正在生成越来越多的海量生物数据集。这些数据的范围从高通量(Meta)基因组测序到社区人口水平的研究,只有通过协调电子健康记录的出现才有可能。研究人员的重点一直放在分析和解释这些数据集,以产生生物学见解或医疗干预。然而,这种对生物影响的关注掩盖了处理和传输这些数据的基本基础设施挑战,为此必须设计适当和有针对性的数据压缩技术。随着计算资源和传输带宽变得无法处理涌入,更快的算法和适当的数据压缩对于大规模分析变得至关重要。幸运的是,与简单地应用通用数据压缩不同,有针对性的压缩方法的设计通常会导致发现其他理想的生物相关特征。该项目的目标是(1)设计适合快速传输和分析基因组测序数据的新的有损压缩特征集,(2)开发用于隐私保护分布式分析的医疗数据的简洁压缩摘要草图,以及(3)利用前两个目标的见解来构建更快的生物分析软件。目标和方法(1)压缩算法通常依赖于源数据中重复模式的识别来构造压缩表示。在测序数据的背景下,生物学家通常依赖于随机k-mer选择来寻找冗余。我们相信,严格分析k-mer选择方法和相关的替代方案,可以用来利用冗余的人口作图和宏基因组数据集。(2)识别重复模式的另一种方法是只提取下游代理感知的模式。在这种模式下,对于许多分析,我们不需要访问原始数据,而是可以使用概率摘要。这不仅有助于减少合作机构之间的传输要求,而且还可以改善和提供隐私保障,这在处理患者健康记录时很有用。这些概率摘要可以进一步用来自密码学文献的多方计算技术来增强,以向参与分析的所有各方提供隐私和安全保证。(3)从以前的工作中,我们知道,在构建较小的数据表示时,我们通常可以提高下游分析算法的运行时间和准确性。这不是一个单独的目标。该提案的一个目的是通过设计和原型化可用的软件包和库,向从业者展示目标(1)和(2)的压缩表示的实际相关性。
英文摘要
Increasingly massive biological data sets are being generated. These data range from high-throughput (meta)genomic sequencing to population-level studies of communities made possible only by the advent of coordinated electronic health records. Much of the focus of researchers has been on the analysis and interpretation of these data sets for generating biological insights or medical interventions. However, this focus on biological impact obscures the underlying fundamental infrastructural challenges of handling and transmitting those data, for which the design of appropriate and targeted data compression techniques is essential. As computational resources and transmission bandwidth become incapable of handling the influx, faster algorithms and appropriate data compression become essential for large scale analytics. Fortunately, unlike simply applying general-purpose data compression, the design of targeted compression methods often leads to the discovery of other desirable biologically-relevant features. The aims of this project are (1) to design new lossy compressive feature sets suitable for fast transmission and analysis of genomic sequencing data, (2) to develop succinct compressed summary sketches of medical data for privacy-preserving distributed analyses, and (3) to utilize insights from the previous two aims to build faster bioanalysis software. GOALS and APPROACH (1) Compression algorithms typically rely on the identification of repetitive patterns in the source data to structure the compressed representation. In the context of sequencing data, biologists have often relied on random k-mer selection to find redundancies. We believe that rigorously analyzing k-mer selection methods and related alternatives for can be used to exploit redundancy in both population mapping and metagenomic data sets. (2) An alternative to identifying repetitive patterns is to extract only patterns that downstream agents perceive. In this mode, for many analyses, we do not need access to the raw data, but can instead work with probabilistic summaries. This not only assists in reducing transmission requirements between collaborating institutions, but can also improve and provide privacy guarantees, useful when dealing with patient health records. These probabilistic summaries can further be augmented with multi-party computation techniques from the cryptographic literature to give privacy and security guarantees to all of the parties involved in the analysis. (3) From prior work, we know that it often turns out that in the building of smaller representations of data, we can often improve the runtime and sometimes accuracy of downstream analysis algorithms. This is not so much a separate goal. One aim of this proposal is to demonstrate the practical relevant of the compressed representations from Goals (1) and (2) to practitioners through the design and prototyping of usable software packages and libraries.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Data compression for biomedical data analysis
-
批准号:DGDND-2022-03074
-
项目类别:DND/NSERC Discovery Grant Supplement
-
资助金额:$2.91万
-
财政年份:2022
-
负责人:Yu, YunWilliam
-
依托单位:
Data compression for biomedical data analysis
-
批准号:DGECR-2022-00353
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2022
-
负责人:Yu, YunWilliam
-
依托单位:
海外基金