课题基金 / 基金详情

项目摘要

项目成果

Nathan Sheffield的其他基金

相似基金

相关文献

中文摘要
翻译
摘要 这一行政补充为表观基因组间隔数据创建了AI/ML就绪的资源。 表观基因组数据总结为基因组间隔集合,现在可用于数千种细胞变异 类型、疾病、状况等。这些数据为理解基因调控和疾病- 原因许多健康结果受到基因变异或调控DNA表观遗传扰动的影响。这个 Parent R01开发了新的、可扩展的算法和基因组区间数据集之间的相似性度量。 这些进展将提高现有生物医学研究方法的效率和准确性,这些方法可以 依靠分析基因组区域数据。它们将为探索广袤和不断增长的 基因组间隔数据语料库。 在这份管理补充资料中,我们寻求利用这个丰富的数据源,并为 社区。虽然已经做出了一些努力来创建基因组间隔数据的统一处理的数据库, 目前很少有高质量的基因组间隔是为机器学习应用设计的。 跨数据源集成表观基因组数据的第一个步骤之一是defifiConsensus Regions,该区域 fit t原始数据。许多下游分析,特别是学习任务,都依赖于这样的共识 区域设置。然而,选择一个好的共识可能是一个既耗时又令人困惑的过程,而且 可能会丢失大量信息并在结果中引入错误。为了帮助缓解这一挑战, 这项提议将通过一种原则性的方法利用几个数据集来生成AI/ML就绪的资源。这 过程将包括1)去fi宁化共识区域;2)将原始数据投射到共识中以使其标准化; 3)规范标注。最后,我们将以用户友好的方式向社区提供这些服务 记录良好的界面。结果将是一系列可供社区使用的数据集 构建ML模型。
英文摘要
ABSTRACT This administrative supplement creates AI/ML-ready resources for epigenome genomic interval data. Epigenome data summarized as sets of genomic intervals are now available for thousands of variations of cell type, disease, condition, etc. This data holds tremendous promise to understand gene regulation and disease be- cause many health outcomes are affected by genetic variation or epigenetic perturbation in regulatory DNA. The parent R01 develops novel, scalable algorithms and measures of similarity between genomic interval datasets. These advances will improve both the efficiency and accuracy of existing biomedical research approaches that rely on analyzing genomic region data. They will open the door to new ways of exploring the vast and growing corpus of genome interval data. In this administrative supplement, we seek to take this rich data source and produce AI/ML-ready resources for the community. While there has been some effort to create uniformly processed databases of genomic interval data, there are few high-quality genomic interval currently available that are designed for machine learning applications. One of the first steps to integrating epigenome data across data sources is defining consensus regions that fit the original data well. Many downstream analyses, particularly learning tasks, rely on such a consensus region set. However, choosing a good consensus can be a time-consuming and confusing process, and also has potential to lose substantial information and introduce errors into results. To help alleviate this challenge, this proposal will take several datasets through a principled approach to generate AI/ML-ready resources. This process will include 1) defining consensus regions; 2) projecting raw data into the consensus to standardize it; and 3) standardizing annotation. Finally, we will make these available to the community with user-friendly and well-documented interfaces. The outcome will be a series of datasets that are ready for use for the community to build ML models.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Novel methods for large-scale genomic interval comparison
  • 批准号:
    10678947
  • 项目类别:
  • 资助金额:
    $38.4万
  • 财政年份:
    2022
  • 负责人:
    Nathan Sheffield
  • 依托单位:
A modular data analysis ecosystem using portable encapsulated projects
  • 批准号:
    10468680
  • 项目类别:
  • 资助金额:
    $39.38万
  • 财政年份:
    2018
  • 负责人:
    Nathan Sheffield
  • 依托单位:
A modular data analysis ecosystem using portable encapsulated projects
  • 批准号:
    10019399
  • 项目类别:
  • 资助金额:
    $39.38万
  • 财政年份:
    2018
  • 负责人:
    Nathan Sheffield
  • 依托单位:
A modular data analysis ecosystem using portable encapsulated projects
  • 批准号:
    9751344
  • 项目类别:
  • 资助金额:
    $39.32万
  • 财政年份:
    2018
  • 负责人:
    Nathan Sheffield
  • 依托单位:
海外基金