III: Small: Topology-based approaches to integrated analysis of transcriptomic, protein interactomic and phenotypic data
III: Small: Topology-based approaches to integrated analysis of transcriptomic, protein interactomic and phenotypic data
批准号:
1218201
负责人:
Jianhua Ruan
金额:
$45.27万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-10-01 至 2017-09-30
中文摘要
高通量技术现在可以同时测量细胞中成千上万个分子的活动和相互作用,为生物学的系统级科学探索打开了新的大门。先进的计算方法正在开发中,用于分析大量数据,提取代表知识的模式,并构建预测模型,这些模型有望确定生物体的特征,例如癌症结果或植物生长表型。为了帮助实现这些目标,该项目旨在开发高效和有效的计算算法和工具,以整合异构和嘈杂的高通量数据,并以将基因视为相互连接而不是细胞独立组成部分的方式进行分析。生物网络是描述细胞中分子之间相互作用的数学模型,对复杂生物系统的建模和理解至关重要。然而,生物网络的庞大规模和复杂性以及数据的噪声和不完整给基于网络的数据分析带来了严峻的挑战。为了应对这些挑战,一个实用和直观的策略是在功能通路水平上分析/利用这些网络,即参与类似生物过程的基因/蛋白质,这将大大降低生物网络的复杂性,提高对复杂表型的理解。由于目前对大多数物种的功能通路的了解相当有限,本项目将开发一套算法和软件工具,用于全自动发现密集子网络作为候选功能模块,并开发面向功能模块的算法,用于分析/利用生物网络的几个实际应用。首先,该项目将开发算法,以提高网络质量和网络模块发现利用嵌入在网络拓扑中的信息。对于网络(例如蛋白质-蛋白质相互作用)数据,利用拓扑来提高边缘可靠性,随后使用基于图上随机行走的新颖拓扑相似性度量来发现模块。对于非网络化(例如,转录组学)数据,利用全局网络拓扑结构构建一个“最优”网络,使完全自动化的模块发现无需任何用户指定的参数。其次,本研究将开发计算方法来系统地研究网络拓扑与生物功能之间的关系,这有望推进目前对生物网络组织原理的理解,并促进疾病研究中基因的优先排序。最后,本项目提出了一种新的基于Steiner树的算法来识别与癌症表型相关的潜在因果基因,以及一种基于途径/子网水平基因表达模式的新的相似性度量来比较患者,该相似性度量可以很容易地与现有的聚类/分类算法相结合,用于基于网络的癌症预后预测。该项目的最终产出将包括用于综合数据分析的生物信息学工具和从不同输入数据集中发现的生物知识数据库。这些工具和资源将在网络上免费提供,可供广泛的对生物信息学算法开发或应用感兴趣的研究人员使用。这些工具和资源将被应用于研究合作者感兴趣的几个生物过程,他们承诺验证一些计算预测。这些研究包括通过整合蛋白质-蛋白质相互作用和转录组学数据,鉴定新的植物激素反应基因,预测和表征DNA损伤反应基因,以及预测乳腺癌患者的转移潜力。该项目还将通过开发新的网络链路预测和模块发现算法以及网络约束聚类/分类方法来促进计算的进步,这些方法有望在生物科学以外的其他领域得到直接应用。作为这项研究的一部分,所开展的活动将被纳入几门课程,并将扩大德克萨斯大学圣安东尼奥分校的教育和研究机会,这是一所为少数族裔服务的学院,大多数本科生来自代表性不足的少数族裔,预计将增加地理和种族多样性,并鼓励少数族裔参与生物信息学和计算生物学研究。
英文摘要
High-throughput technology now allows measuring the activities and interactions of tens of thousands of molecules in the cell simultaneously, opening new doors to systems-level scientific exploration in biology. Advanced computational methods are in development to analyze the huge amount of data to extract patterns that represent knowledge and to construct predictive models that have the promise to determine the characteristics of an organism, such as cancer outcomes or plant growth phenotypes. To help achieve these goals, this project aims at developing efficient and effective computational algorithms and tools to integrate heterogeneous and noisy high-throughput data, and to analyze them in ways that treat genes as inter-connected rather than independent components of the cell. Biological networks are mathematical models that describe the interactions among molecules in the cell and are critical to the modeling and understanding of complex biological systems. However, the large sizes and complexity of biological networks as well as the noisy and incomplete data pose critical challenges to network-based data analysis. To tackle these challenges, a practical and intuitive strategy is to analyze/utilize such networks on the level of functional pathways, i.e., genes/proteins involved in similar biological processes, which would significantly reduce the complexity of biological networks and improve the understanding of complex phenotypes. As the current knowledge of functional pathways is rather limited for most species, this project will develop a set of algorithms and software tools for fully automated discovery of dense subnetworks as candidate functional modules, and develop functional module-oriented algorithms for analyzing/utilizing biological networks for several real applications. First, this project will develop algorithms to improve network quality and network module discovery using information embedded in network topology. For networked (e.g. protein-protein interaction) data, topology is utilized to improve edge reliability, and subsequently module discovery, using a novel topological similarity measurement based on random walks on graphs. For non-networked (e.g., transcriptomic) data, global network topology is utilized to construct an "optimal" network that enables fully automated module discovery without any user-specified parameters. Second, this research will develop computational methods to systematically investigate the relationship between network topology and biological functions, which is expected to advance the current understanding of the organizing principles of biological networks, and facilitate prioritizing genes in disease studies. Finally, this project proposes a novel Steiner tree based algorithm for identifying potential causal genes associated with cancer phenotypes, and a novel similarity metric to compare patients based on pathway/subnetwork-level gene expression patterns, which can be easily combined with existing clustering/classification algorithms for network-based prediction of cancer outcomes. The final outputs of this project will include both bioinformatics tools for integrative data analysis and databases of biological knowledge discovered from different input datasets. These tools and resources will be made freely available on the web, which can be used by a broad range of researchers who are interested in bioinformatics algorithm development or applications. These tools and resources will be applied to study several biological processes of central interests to collaborators, who have committed to validate some of the computational predictions. These include identifying novel plant hormone response genes, predicting and characterizing DNA damage response genes, and predicting metastasis potentials for breast cancer patients, by integrating protein-protein interaction and transcriptomic data. This project will also contribute to the advancement of computing with the development of novel network link prediction and module discovery algorithms and network-constrained clustering/classification methods that are expected to have immediate applications in other domains besides biological sciences. The activities undertaken as part of this research will be incorporated into several courses and will expand the educational and research opportunities available at the University of Texas at San Antonio, a minority-serving institute where the majority of undergraduates are from under-represented minorities, and is expected to increase the geographic and ethnic diversity and encourage the participation of minority groups in bioinformatics and computational biology research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
ABI Innovation: Tools and databases for network-based plant systems biology with applications to understanding plant-virus interactions
-
批准号:1565076
-
项目类别:Standard Grant
-
资助金额:$68.38万
-
财政年份:2016
-
负责人:Jianhua Ruan
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: