课题基金 / 基金详情

CAREER: Large-scale biological network integration with applications to automated function annotation

CAREER: Large-scale biological network integration with applications to automated function annotation
职业:大规模生物网络集成与自动化功能注释的应用
批准号:
1652815
负责人:
Jian Peng
金额:
$78.32万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-04-15 至 2022-03-31

项目摘要

项目成果

Jian Peng的其他基金

相似基金

相关文献

中文摘要
翻译
该项目旨在开发新的方法,用于整合来自不同类型分子和测量方法的大量高分辨率数据;目标是确定分子如何随着时间的推移相互作用,以实现基本的生物功能。生物功能通过生物分子的无数相互作用来实现,例如当蛋白质结合到其他蛋白质以调节其活性或核酸以调节基因时。DNA测序导致了报告基因组序列及其变异的数据的爆炸性增长,以及通过转录本分析的基因表达;用于蛋白质和代谢物分析的高通量技术的数据流正在迅速赶上。这导致了不断扩大的存储库,这些存储库可以存档,组织和共享所产生的数据:通过将实验条件与分子特征相联系,研究人员可以了解发生了哪些分子相互作用,并从中推断出活细胞中的许多生物学功能。从这些数据集中提取有意义的生物学见解在两个方面具有挑战性:数据集非常大,因此它们需要计算方法进行基本处理,并且每种类型的数据在许多方面都与其他数据不同(噪声类型,错误源,完整性等)。因此它们可能需要不同的统计建模,以便在合并它们之前正确地标准化它们。如果正确执行,所得的高维数据集适用于各种预测分析,揭示分子相互作用组中的功能模块。该项目的成果将通过网络服务器和开放源码软件提供。综合研究和教育活动包括跨学科的生物信息学课程开发,推广到高中学生和研究机会,为学生在代表性不足的群体。全面了解基因或蛋白质的各种功能方面,如参与特定的生物过程,物理/遗传相互作用,或疾病关联,是生物学和转化医学研究的关键。由于通过生物实验详尽地表征基因或蛋白质通常是棘手的,因此知识和计算假设生成的系统级集成作为指导实验的有效方法在该领域引起了极大的兴趣。在这个项目中,我们将开发一个新的计算框架,用于异构网络和功能基因组数据的数据集成和降维,以获得低维向量空间中信息丰富的数据表示。为了利用分子网络和进化信息,我们将应用所提出的降维技术来有效地整合多个物种的序列数据和网络数据,以预测基因功能。 我们的方法将使大规模的,综合的,跨物种的,基因组规模的基因功能注释。通过这种整合,我们的方法也可以推断功能同源性或基因之间的相似性,共享弱序列相似性,但相关的生物学功能,从不同的物种。结果、软件和其他信息将在http://jianpeng.cs.illinois.edu上提供。
英文摘要
This project aims to develop new methods for integrating large amounts of high resolution data arising from different types of molecules and measurement methods; the goal is to ascertain how the molecules interact over time to carry out essential biological functions. Biological functions are carried out through the myriad interactions of biological molecules, such as when proteins bind to other proteins to modulate their activity or to nucleic acids to regulate genes. DNA sequencing has led to an explosive growth in data reporting genomic sequences and their variations, and gene expression through transcript profiling; the data streams from high-throughput technologies for protein and metabolite profiling are quickly catching up. This has led to ever-expanding repositories that archive, organize and share the resulting data: by also connecting experimental conditions to the molecular profiles, researchers come to understand which molecular interactions occur and, from these, deduce many of the biological functions in living cells. Extraction of meaningful biological insights from these data sets is challenging in two ways: the data sets are very large so they require computational methods for basic handling, and each type of data differs from the others in many ways (type of noise, source of error, completeness, etc.) so they may require different statistical modeling to standardize them correctly prior to merging them. Carried out correctly, the resulting high-dimensional data sets are suitable for a variety of predictive analytics that reveal functional modules in the molecular interactomes. Results from this project will be made available through webservers and open source software. The integrated research and educational activities include interdisciplinary bioinformatics curriculum development, outreach to high school students and research opportunities for students in underrepresented groups.Comprehensively understanding various functional aspects of a gene or a protein, such as involvement in a particular biological process, physical/genetic interactions, or disease association, is critical for both biology and translational medicine research. Since exhaustively characterizing genes or proteins through biological experiments is often intractable, systems-level integration of knowledge and computational hypothesis generation have garnered great interest in the field as an effective way to guide experiments. In this project, we will develop a novel computational framework for data integration and dimensionality reduction of heterogeneous network and functional genomic data to obtain informative data representations in a low-dimensional vector space. To utilize both molecular networks and evolutionary information, we will apply the proposed dimensionality reduction techniques to effectively integrate sequence data and network data across multiple species for predicting gene function. Our approaches will enable large-scale, integrated, cross-species, genome-scale gene function annotation. Through this integration, our methods can also infer functional homology or analogy between genes, which share weak sequence similarity but relevant biological functions, from different species. Results, software and additional information will be available at http://jianpeng.cs.illinois.edu.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
Large-scale integration of heterogeneous pharmacogenomic data for identifying drug mechanism of action
大规模整合异质药物基因组数据以识别药物作用机制
DOI: 10.1142/9789813235533_0005
发表时间: 2018
期刊: Proceedings of the Pacific Symposium
影响因子: --
作者: [Luo, Yunan, Wang, Sheng, Xiao, Jinfeng, Peng, Jian]
通讯作者: Peng, Jian
DOI: 10.1371/journal.pcbi.1007283
发表时间: 2019-09
期刊: PLoS Computational Biology
影响因子: 4.3
作者: [Yufeng Su;Yunan Luo;Xiaoming Zhao;Yang Liu;Jian Peng]
通讯作者: Yufeng Su;Yunan Luo;Xiaoming Zhao;Yang Liu;Jian Peng
Framework: Software: NSCI: Collaborative Research: Hermes: Extending the HDF Library to Support Intelligent I/O Buffering for Deep Memory and Storage Hierarchy Systems
国内基金
海外基金
基于水稻穗粒数关键基因LARGE2提高作物产量的探索与应用
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    黄洛将
  • 依托单位:
水稻穗粒数调控关键因子LARGE6的分子遗传网络解析
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    黄洛将
  • 依托单位:
量子自旋液体中拓扑拟粒子的性质:量子蒙特卡罗和新的large-N理论
  • 批准号:
    12074246
  • 项目类别:
    面上项目
  • 资助金额:
    62.0万元
  • 批准年份:
    2020
  • 负责人:
    Yoshitomo Kamiya
  • 依托单位:
甘蓝型油菜Large Grain基因调控粒重的分子机制研究
  • 批准号:
    31972875
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    石江华
  • 依托单位: