课题基金 / 基金详情

Data compression for biomedical data analysis

Data compression for biomedical data analysis
用于生物医学数据分析的数据压缩
批准号:
RGPIN-2022-03074
负责人:
Yu, YunWilliam
金额:
$2.11万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Yu, YunWilliam的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Increasingly massive biological data sets are being generated. These data range from high-throughput (meta)genomic sequencing to population-level studies of communities made possible only by the advent of coordinated electronic health records. Much of the focus of researchers has been on the analysis and interpretation of these data sets for generating biological insights or medical interventions. However, this focus on biological impact obscures the underlying fundamental infrastructural challenges of handling and transmitting those data, for which the design of appropriate and targeted data compression techniques is essential. As computational resources and transmission bandwidth become incapable of handling the influx, faster algorithms and appropriate data compression become essential for large scale analytics. Fortunately, unlike simply applying general-purpose data compression, the design of targeted compression methods often leads to the discovery of other desirable biologically-relevant features. The aims of this project are (1) to design new lossy compressive feature sets suitable for fast transmission and analysis of genomic sequencing data, (2) to develop succinct compressed summary sketches of medical data for privacy-preserving distributed analyses, and (3) to utilize insights from the previous two aims to build faster bioanalysis software. GOALS and APPROACH (1) Compression algorithms typically rely on the identification of repetitive patterns in the source data to structure the compressed representation. In the context of sequencing data, biologists have often relied on random k-mer selection to find redundancies. We believe that rigorously analyzing k-mer selection methods and related alternatives for can be used to exploit redundancy in both population mapping and metagenomic data sets. (2) An alternative to identifying repetitive patterns is to extract only patterns that downstream agents perceive. In this mode, for many analyses, we do not need access to the raw data, but can instead work with probabilistic summaries. This not only assists in reducing transmission requirements between collaborating institutions, but can also improve and provide privacy guarantees, useful when dealing with patient health records. These probabilistic summaries can further be augmented with multi-party computation techniques from the cryptographic literature to give privacy and security guarantees to all of the parties involved in the analysis. (3) From prior work, we know that it often turns out that in the building of smaller representations of data, we can often improve the runtime and sometimes accuracy of downstream analysis algorithms. This is not so much a separate goal. One aim of this proposal is to demonstrate the practical relevant of the compressed representations from Goals (1) and (2) to practitioners through the design and prototyping of usable software packages and libraries.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Data compression for biomedical data analysis
  • 批准号:
    DGDND-2022-03074
  • 项目类别:
    DND/NSERC Discovery Grant Supplement
  • 资助金额:
    $2.91万
  • 财政年份:
    2022
  • 负责人:
    Yu, YunWilliam
  • 依托单位:
Data compression for biomedical data analysis
  • 批准号:
    DGECR-2022-00353
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2022
  • 负责人:
    Yu, YunWilliam
  • 依托单位:
海外基金