IIBR Informatics: Mixture model algorithms for inferring covariance structures and microbial associations from microbiome data
IIBR Informatics: Mixture model algorithms for inferring covariance structures and microbial associations from microbiome data
批准号:
2400009
负责人:
Shibu Yooseph
金额:
$63.47万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
已结题
起止时间:
2023-10-01 至 2024-05-31
中文摘要
微生物群落在地球上几乎无处不在,它们在它们所处的环境中发挥着重要的功能作用。群落中的微生物在争夺环境中可用的食物和能源资源时相互作用。微生物之间的这些直接和间接的相互作用,称为微生物协会,在决定群落的结构、组织和功能方面发挥着重要作用。该项目解决了从使用高通量DNA测序技术产生的微生物组数据推断微生物相关性的计算挑战。该项目开发的新的计算工具和资源将促进几个学科的知识进步,包括环境科学、医学和人类健康科学。该项目将有助于理解微生物生态系统的生命规律,并将进一步加深我们对微生物在环境中的生物地球化学过程和与微生物相关的疾病的发展中所发挥的重要作用的理解。该项目将为研究生提供跨学科培训,重点是培训代表性不足的群体(包括妇女和少数群体)。该项目还将通过开发一个教育模块,通过讲习班向高中教师介绍基因组学和生物信息学的介绍性主题,从而有助于提高高中生在STEM领域的参与度。从微生物类群丰度所确定的潜在协方差结构中可以推断出微生物的关联性。这些丰度通常是根据DNA序列数据估计的。然而,序列数据本质上是组成的,因为它们只提供类群的相对丰富的信息,这给确定微生物相关性带来了挑战。此外,微生物类群之间的关联并不总是固定的,当资源可获得性和环境特征等因素发生变化时,它们可能会发生变化。该项目将开发新的计算方法,以确定大型微生物组数据集中协方差结构的数量,并重建微生物关联组。这些方法将能够从序列数据中捕获积极和消极的微生物关联,同时处理序列数据的组成性质带来的挑战。总体方法是以混合模型框架为基础的,该框架结合了对微生物丰度数据进行建模的成分分布。该项目将开发变分近似算法来确定给定微生物组数据集中的协方差结构的数量,开发快速数值优化算法来估计混合模型的参数,并开发一个综合框架来将元数据纳入分析。这些算法还将能够重建稀疏模型,从而处理群落中微生物协会数量较少的情况。将这些算法应用于分析大型微生物组数据集,将对三种不同环境(人类、海洋和土壤)的微生物生态产生新的见解。这一分析将包括阐明菌株水平上的微生物组合、潜在微生物网络的结构以及这些环境中关键类群的身份。该项目的结果可以在https://github.com/syooseph/YoosephLab/tree/master/MixtureMicrobialNetworks.This上找到,该奖项反映了国家科学基金会的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Microbial communities are found almost everywhere on earth and they play important functional roles in the environments that they are found in. Microbes in a community interact with each other as they compete for the food and energy resources available in their environment. These direct and indirect interactions between microbes, termed microbial associations, play a large role in determining the structure, organization, and function of the community. This project addresses the computational challenge of inferring microbial associations from microbiome data generated using high-throughput DNA sequencing technologies. The novel computational tools and resources developed by this project will enable the advancement of knowledge in several disciplines, including environmental sciences, medicine, and human health science. This project will contribute to understanding the rules of life for microbial ecosystems, and it will further our understanding of the important roles that microbes play in biogeochemical processes in the environment and in the progression of microbe-associated diseases. This project will provide interdisciplinary training for graduate students, with an emphasis on training under-represented groups (including women and minorities). This project will also contribute to enabling an increased level of high school student participation in STEM areas through the development of an education module that will introduce high-school teachers, via workshops, to introductory topics in genomics and bioinformatics. Microbial associations can be inferred from the underlying covariance structure that is determined from microbial taxa abundances. These abundances are often estimated from DNA sequence data. However, sequence data are compositional in nature, in the sense that they only provide relative abundance information for taxa, and this poses challenges when determining microbial associations. Furthermore, associations between groups of microbial taxa are not always fixed, and they can change when factors such as resource availability and environmental characteristics vary. This project will develop novel computational methods to determine the number of covariance structures in large microbiome datasets and to reconstruct the sets of microbial associations. These methods will be able to capture both positive and negative microbial associations from sequence data while dealing with the challenges posed by the compositional nature of sequence data. The overall approach is based on a mixture model framework incorporating component distributions that model microbial abundance data. This project will develop variational approximation algorithms to determine the number of covariance structures in a given microbiome dataset, fast numerical optimization algorithms to estimate the parameters of the mixture model, and an integrated framework to incorporate metadata in the analysis. The algorithms will also enable the reconstruction of sparse models, thus handling the scenario when the number of microbial associations in the community is small. The application of these algorithms to analyze large microbiome datasets will generate new insights into microbial ecology of three different environments (human, ocean, and soil). This analysis will include an elucidation of microbial associations at the strain level, the structures of the underlying microbial networks, and the identities of the key taxa in these environments. The results of the project can be found at https://github.com/syooseph/YoosephLab/tree/master/MixtureMicrobialNetworks.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
IIBR Informatics: Mixture model algorithms for inferring covariance structures and microbial associations from microbiome data
-
批准号:2051283
-
项目类别:Standard Grant
-
资助金额:$63.47万
-
财政年份:2021
-
负责人:Shibu Yooseph
-
依托单位:
ABI Development: A Novel Protein Fragment Assembler for Metagenomic Data Analysis
-
批准号:1262295
-
项目类别:Continuing Grant
-
资助金额:$151.63万
-
财政年份:2013
-
负责人:Shibu Yooseph
-
依托单位:
海外基金