课题基金 / 基金详情

BIGDATA: DA: Interpreting massive genomic data sets via summarization

BIGDATA: DA: Interpreting massive genomic data sets via summarization
BIGDATA:DA:通过汇总解释海量基因组数据集
批准号:
8840551
负责人:
William Stafford Noble
金额:
$21.9万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-04-01 至 2017-03-31

项目摘要

项目成果

William Stafford Noble的其他基金

相似基金

相关文献

中文摘要
翻译
基因组数据很大,而且还在变得越来越大,但目前的分析方法还不能扩展到对数千或数百万基因组的分析。因此,一个关键的技术挑战是开发能够分析这些巨大数据集的新方法。在这个建议中,我们描述了一个新的计算框架,用于从海量基因组数据集中提取推理。我们的方法利用了已经开发的用于分析文本语料库的子模块摘要方法。我们将把这些方法应用于基因组学中的五个大数据问题:1)识别特定人类细胞类型的功能元件;2)识别与特定癌症亚类相关的基因组特征;3-4)识别代表祖先或表型定义的人类群体的基因组变体;以及5)找到一组表征人体上特定位置的微生物基因。该项目将在两个方面促进发现和理解。首先,我们将开发新的方法来总结基因组、表观基因组和元基因组数据集。事实上,据我们所知,这项拨款提出了首次将摘要方法应用于任何类型的基因组数据。拟议的研究将大大提高我们将子模块应用于这些汇总任务的能力,特别是在确定和创建已针对提案中概述的五项任务进行验证的距离函数库方面。第二,我们将把我们的新方法运用到具有深远意义的问题上。事实上,在我们五项任务中的任何一项上取得重大进展,都将代表着我们对人类历史、生物学或疾病的科学理解取得了重要进展。该项目的影响将随着大数据问题的增长而增长,即使在项目完成后也是如此。这个项目的结果,无论是我们开发的软件还是我们制作的摘要,都将有助于回答任何必须处理大数据的领域的广泛问题。
英文摘要
Genomic data is big and getting ever bigger, but current analysis methods will not scale to the analysis of thousands or millions of genomes. Consequently, a critical technical challenge is to develop new methods that can analyze these enormous data sets. In this proposal, we describe a new computational framework for drawing inferences from massive genomic data sets. Our approach leverages submodular summarization methods that have been developed for analyzing text corpora. We will apply these methods to five big data problems in genomics: 1) identifying functional elements characteristic o f a given human cell type; 2) identifying genomic features associated with a particular subclass of cancer; 3-4) identifying genomic variants representative of ancestrally or phenotypically defined human populations; and 5) finding a set of microbial genes that characterize a given site on the human body. This project will advance discovery and understanding on two fronts. First, we will develop novel methods for summarizing genomic, epigenomic and metagenomic data sets. Indeed, to our knowledge, this grant proposes the first application of summarization methods to genomic data of any kind. The proposed research will significantly advance our ability to apply submodularity to these summarization tasks, particularly with respect to identifying and creating a library of distance functions that have bee validated with respect to the five tasks outlined in the proposal. Second, we will apply our novel methods to problems of profound importance. Indeed, significant progress toward any one of our five tasks would represent an important advance in our scientific understanding of human history, biology or disease. The impact of this project will grow as the big data problem grows, even after the project is complete. The results of this project, both the software that we develop and the summaries that we produce, will be useful for answering a wide array of questions in any field that must cope with big data.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1038/nrg3920
发表时间: 2015-06
期刊: Nature reviews. Genetics
影响因子: --
作者: []
通讯作者:
Deep tensor genomic imputation
  • 批准号:
    10557916
  • 项目类别:
  • 资助金额:
    $38.38万
  • 财政年份:
    2021
  • 负责人:
    William Stafford Noble
  • 依托单位:
Deep tensor genomic imputation
  • 批准号:
    10096947
  • 项目类别:
  • 资助金额:
    $39.86万
  • 财政年份:
    2021
  • 负责人:
    William Stafford Noble
  • 依托单位:
Optimization and joint modeling for peptide detection by tandem mass spectrometry
  • 批准号:
    9214942
  • 项目类别:
  • 资助金额:
    $33.23万
  • 财政年份:
    2017
  • 负责人:
    William Stafford Noble
  • 依托单位:
Project 2: UW-CNOF Data Analysis and Modeling
  • 批准号:
    9021413
  • 项目类别:
  • 资助金额:
    $63.28万
  • 财政年份:
    2015
  • 负责人:
    William Stafford Noble
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis