课题基金 / 基金详情

AF: Small: Redundancy exploiting algorithms for high throughput genomics

AF: Small: Redundancy exploiting algorithms for high throughput genomics
AF:小:利用冗余算法实现高通量基因组学
批准号:
1619081
负责人:
Qin Zhang
金额:
$40.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-08-01 至 2020-07-31

项目摘要

项目成果

Qin Zhang的其他基金

相似基金

相关文献

中文摘要
翻译
确定个体的基因组组成对于理解某些基因组变异最终如何导致疾病(如癌症)至关重要。确定农业上重要的植物、树木、农场动物和野生动物的基因组组成有助于改善农业、林业、兽医和环境科学。自2008年推出“下一代测序技术”以来,基因组测序的成本下降了1000倍。这导致了基因组数据生成速度的提高,远远超过了我们计算和数据存储能力的提高。随着这些廉价、快速的基因组测序技术的出现,科学界已经能够启动大型项目,如全基因组泛癌症分析项目,该项目旨在确定数千名癌症患者的基因组序列。我们的项目旨在通过新的基因组数据压缩方法来解决这些大规模基因组研究中迫在眉睫的数据大小挑战,这些方法旨在减少基因组序列表示方式的冗余。这种冗余的来源是个体患者的基因组序列之间的高度相似性,以及单个人类基因组的基因组区域之间的高度相似性。由于从基因组序列中提取信息的主要困难是计算,通过压缩方法减少管理和分析基因组数据所需的计算资源将有助于基因组学改善人类生活和环境。该项目对学生和人员培训的影响将体现在印第安纳州大学的两门新的研究生课程上:PI Sahinalp教授的基因组数据的数据管理、访问和处理课程,以及PI Ergun教授的压缩算法课程,重点是基因组数据,强调新的大数据范式压缩的影响。这两门课程将融入CS博士课程,以及现有的生物信息学和数据科学硕士课程;它们也旨在吸引更好奇的本科生。核酸测序技术的快速发展重塑了生命科学的几乎每个领域,从农业到生物能源,从环境科学到生物医学。大规模的基因组计划正在从数千名患者或通过收集环境样本的移动的传感器产生拍字节级的数据。随着技术的进步,大多数去医院就诊的人最终都会进行(可能是组织特异性的)基因组测序。将从数千至数百万非模式生物及其种群中收集基因组数据,以评估相应生态系统内的生物多样性。复杂的微生物群落将从数千个地理位置取样,以研究环境条件的影响。此外,这些研究将涉及持续的数据收集工作,目的是通过使用全基因组或全转录组测序来监测生物系统的动态变化。因此,基因组数据的产生将以前所未有的速度发生,需要开发新的算法来帮助减少基因组序列数据对计算、存储和传输系统的负担。该项目结合了印第安纳州大学两位研究人员的独特优势,为基因组学中的关键基础设施问题带来了一种有原则的算法方法。该项目将通过使用新的压缩工具和压缩数据结构来通信、存储、管理和访问大量(流式)基因组数据,来满足大型癌症项目、收集环境样本的便携式设备以及嵌入人体的更小传感器的下一阶段基因组数据生成的需求。为此,我们将采用和扩展现有的算法剧目,涉及近似算法,次线性算法,无损数据压缩,I/O效率,内存层次结构感知/遗忘和压缩数据结构。
英文摘要
Determining the genomic makeup of individuals is crucial for understanding how certain genomic variants ultimately lead to disease (such as cancer). Determining genomic makeup of agriculturally important plants, trees, farm animals and wild life help improve agriculture, forestry, veterinary medicine and environmental science. Since the introduction of "next generation sequencing technologies" in 2008, the cost of genome sequencing has dropped by a factor of 1000. This has led to an increase in the speed genomic data is generated that far outpaces the improvements in our computing and data storage capability. With the advent of these cheap, and fast genome sequencing technologies, the scientific community has been able to launch mega-projects such as The Pan Cancer Analysis of Whole Genomes Project, which aim to determine the genome sequences of thousands of cancer patients. Our project aims to address the imminent data size challenges in these large scale genomic studies through new genomic data compression methods that aim to reduce the redundancy in how genomic sequences are represented. The source of this redundancy is the high similarity among genome sequences of individual patients, as well as the high similarity between regions across the genome of a single human genome. Since the main difficulty in extracting information from genome sequences is computational, reduction in the computational resources needed to manage and analyze genomic data through the compression methods will help genomics improve human life and the environment. The impact of this project on student and personnel training will be in terms of two new graduate courses at Indiana University: a course on data management, access and processing for genomic data by PI Sahinalp, and a course on compressed algorithms with a focus on genomic data, emphasizing the effects of new big data paradigms compression, by PI Ergun. Both courses will fit into the CS PhD program, as well as into the existing Bioinformatics and Data Science Master's programs; they are also intended to attract the more curious undergraduates.The rapid advancement of nucleic acid sequencing technology has re-shaped almost every field of life science, from agriculture to bioenergy, and from environmental science to biomedicine. Large-scale genome projects are producing petabyte-scale data from thousands of patients or by mobile sensors collecting environmental samples. As the technology marches forward, most people who visit hospitals will eventually have their (possibly tissue-specific) genomes sequenced. Genomic data will be collected from thousands to millions of non-model organisms and their populations in order to assess the biodiversity within the corresponding ecosystem. Complex microbial communities will be sampled from thousands of geographic locations to study the influence of environmental conditions. Furthermore, these studies will involve continuous data collection efforts, for the purpose of monitoring the dynamic changes in biosystems by the use of genome-wide or transcriptome-wide sequencing. As a result, genomic data generation is to occur at an unprecedented pace, necessitating the development of novel algorithms to help reduce the burden of genomic sequence data on computational, storage and transmission systems. This project combines the unique strengths of the two investigators at Indiana University, bringing a principled, algorithmic approach to critical infrastructure problems in genomics. The project will address the needs of the next stage of genomic data generation by mega cancer projects, portable devices collecting environmental samples, and even smaller sensors to be embedded in the human body, through the use of new compression tools and compressed data structures for communicating, storing, managing, and accessing large collections of (streaming) genome data. For this purpose, we will employ and expand the existing algorithmic repertoire involving approximation algorithms, sublinear algorithms, lossless data compression, I/O efficient, memory hierarchy aware/oblivious and compressed data structures.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: AF: Small: Parallel Reinforcement Learning with Communication and Adaptivity Constraints
  • 批准号:
    2006591
  • 项目类别:
    Standard Grant
  • 资助金额:
    $24.22万
  • 财政年份:
    2020
  • 负责人:
    Qin Zhang
  • 依托单位:
CAREER:Foundation of Communication-Efficient Distributed Computation and Monitoring
  • 批准号:
    1844234
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $49.97万
  • 财政年份:
    2019
  • 负责人:
    Qin Zhang
  • 依托单位:
BIGDATA: Collaborative Research: F: Efficient Distributed Computation of Large-Scale Graph Problems in Epidemiology and Contagion Dynamics
  • 批准号:
    1633215
  • 项目类别:
    Standard Grant
  • 资助金额:
    $53.01万
  • 财政年份:
    2016
  • 负责人:
    Qin Zhang
  • 依托单位:
AF: Small: Efficient Algorithms for Querying Noisy Distributed/Streaming Datasets
  • 批准号:
    1525024
  • 项目类别:
    Standard Grant
  • 资助金额:
    $44.43万
  • 财政年份:
    2015
  • 负责人:
    Qin Zhang
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: