课题基金 / 基金详情

CAREER: Assigning comprehensive, standardized sample annotations to enhance the ability to discover, use, and interpret millions of –omics profiles

CAREER: Assigning comprehensive, standardized sample annotations to enhance the ability to discover, use, and interpret millions of –omics profiles
职业:分配全面、标准化的样本注释,以增强发现、使用和解释数百万个组学概况的能力
批准号:
2045651
负责人:
Arjun Krishnan
金额:
$70.49万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-05-01 至 2023-05-31

项目摘要

项目成果

Arjun Krishnan的其他基金

相似基金

相关文献

中文摘要
翻译
现在在开放的在线数据库中有近200万个实验,在每个实验中,科学家都记录了特定生物样本中数万个DNA元素、基因或蛋白质的活动数据。这些样本来自人类或其他生物在许多条件下的各种组织,是其他科学家重新分析和发现新生物学的宝贵资源。然而,这些数据严重未被充分利用,因为从所有样本的海洋中找到自己感兴趣的样本仍然非常困难。这是因为大多数样本描述是用难以明确搜索的语言编写的,或者在报告样本源的每个关键方面是不完整的。该项目旨在通过对来自6个物种(人类和5个动物模型)的公开样本进行大规模注释,使数据驱动的生物学民主化,这将使研究人员能够发现相关的已发表数据以进行进一步分析。将开发一个网络界面,为研究人员提供一个访问所有这些注释和相关搜索工具的单一访问点。该项目的见解和工具将对生物学研究人员如何提交、存储、管理、访问、重用和重新共享数据产生深远的影响。提高数据的可发现性和可重用性将提高生物学研究的准确性、效率和可重复性,通过加速科学发现节省资源。本项目将开展的教育/培训活动将导致现代生物信息学“隐藏课程”的正式化和公开传播:在快速变化的生物信息学环境中,抽象的经验技能对于进行大数据分析和研究的整体实践能力至关重要。这一努力将在应用计算研究生物学的专业培训中创造开放性和公平性。该项目将开发新的机器学习方法,使用文本和分子数据为公开可用的组学样本分配全面、标准化的注释。完整和结构化样本描述符的障碍是双重的:样本通常使用非标准的、以非结构化自然语言编写的各种术语来描述,即使是基本属性,如组织或环境,如果它们不是原始研究中考虑的因素,也会从样本描述中省略。该项目的目标是消除这两个障碍,并将整合最先进的机器学习(ML)进展,以:开发ML方法,从纯文本描述中推断标准化注释,共同从多个组学类型;开发机器学习模型,从多个物种的分子组学档案中预测结构化元数据;并开发方法,将这些基于文本和组学的模型整合起来,对数百万个样本进行全面注释,并为研究人员提供工具,利用这些巨大的资源来收集新的生物学。与这项研究相结合的是一项教育计划,旨在发展与数据驱动生物学中隐藏课程的正规化和公开传播的联系。该项目的所有结果,包括数据和代码将在https://www.thekrishnanlab.org.This上提供,该奖项反映了NSF的法定使命,并通过使用基金会的智力价值和更广泛的影响审查标准进行评估,认为值得支持。
英文摘要
There are nearly 2 million experiments now in open online databases where, in each experiment, scientists have recorded data on the activity of tens of thousands of DNA elements, genes, or proteins in a particular biological sample. These samples come from a variety of tissues in human or other organisms under numerous conditions and are an invaluable resource for other scientists to reanalyze and discover new biology. However, these data are severely underused because finding the samples one is interested in from the sea of all samples is still very hard. This is because most sample descriptions are written using language that is hard to search unambiguously or are incomplete in reporting every critical aspect of the sample source. This project seeks to democratize data-driven biology by annotating publicly available samples from six species (human and five animal models) on a massive scale, which will enable researchers to discover relevant published data for further analysis. A web-interface will be developed to provide researchers a single access point to all these annotations and the related search tools. The insights and tools from this project will have far-reaching implications for how the biological researchers will submit, store, manage, access, reuse, and re-share data. Improving data discoverability and reusability will increase the accuracy, efficiency, and reproducibility of biological research overall, saving resources by accelerating scientific discovery. The educational/training activities that will be developed in this project will result in formalizing and openly disseminating the “hidden curriculum” in modern bioinformatics: the abstract experiential skills critical for holistic, practical competency in conducting large data analysis and research in a rapidly changing bioinformatics landscape. This effort will create openness and equity in professional training in applying computing to study biology.This project will develop new machine learning methods that use both text and molecular data to assign comprehensive, standardized annotations to publicly-available omics samples. The barrier for complete and structured sample descriptors is two-fold: Samples are routinely described using non-standard, varied terminologies written in unstructured natural language, and Even basic attributes, e.g., tissue or environment, are omitted from sample descriptions if they were not factors considered in the original study. The objective of this project is to remove both of these barriers and will integrate state-of-the-art machine learning (ML) advances to: develop ML methods to infer standardized annotations from plain text descriptions, jointly from multiple omics types; develop ML models to predict structured metadata from molecular omics profiles, jointly from multiple species; and develop methods to integrate these text- and omics-profile-based models to comprehensively annotate millions of samples, and tools for researchers to use this massive resource to glean novel biology. Integrated with this research is an education plan to develop ties to formalize and openly disseminate the hidden curriculum in data-driven biology. All the results from this project including data and code will be available at https://www.thekrishnanlab.org.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: RESEARCH-PGR: Predicting Phenotype from Molecular Profiles with Deep Learning: Topological Data Analysis to Address a Grand Challenge in the Plant Sciences
  • 批准号:
    2310357
  • 项目类别:
    Standard Grant
  • 资助金额:
    $33.39万
  • 财政年份:
    2023
  • 负责人:
    Arjun Krishnan
  • 依托单位:
CAREER: Assigning comprehensive, standardized sample annotations to enhance the ability to discover, use, and interpret millions of –omics profiles
  • 批准号:
    2328140
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $70.49万
  • 财政年份:
    2022
  • 负责人:
    Arjun Krishnan
  • 依托单位:
First Passage Percolation and Related Models
  • 批准号:
    2002388
  • 项目类别:
    Standard Grant
  • 资助金额:
    $2.0万
  • 财政年份:
    2020
  • 负责人:
    Arjun Krishnan
  • 依托单位:
海外基金