CAREER: Scalable binning algorithms for genome-resolved metagenomics
CAREER: Scalable binning algorithms for genome-resolved metagenomics
批准号:
1845890
负责人:
Jason Kwan
金额:
$118.86万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-07-01 至 2024-06-30
中文摘要
地球上几乎每一个环境都是微生物群落的家园,它们的新陈代谢深刻地影响着所有其他生命。因此,了解这样的群落对许多领域都很重要,如农业、生物地球化学、海洋学、生物学、生态学等。然而,绝大多数环境微生物从未在实验室中分离和培养过。因此,了解微生物群落的一个关键是通过培养独立测序(宏基因组学)来解析基因组,其中单个物种的基因组是在一个称为“分组”的过程中从混合宏基因组中重建的。然而,不准确的分类妨碍了对微生物群落获得基本见解的能力,例如在不确定是否从数据集中恢复了所有特定基因组的情况下,或者在一个分类箱中的所有序列确实来自同一基因组的情况下。在这方面,现有的分类方法存在各种问题,包括:(1)在高度复杂的宏基因组中表现不佳;(2)假设所有输入序列都是细菌,导致不纯的分类箱;(3)缺乏对分类箱具有生物学意义的内部验证;(4)忽略了物种内存在多个密切相关的菌株和基因组变异(“泛基因组”概念)。该项目的目标是开发一种可以处理地球上最多样化的微生物群落(如土壤)的新颖性和复杂性的分类方法。所开发的方法有可能改变微生物群落的研究,即使在与基因组未知的高等生物相关的高度复杂的群落中,也能确定“谁在做什么?”随着通过基因组解析宏基因组学对微生物群落的了解的增加,最终将有可能模拟和预测群落的复杂紧急行为,并描绘它们的能力与孤立的组成生物有何不同。该项目还旨在解决目前美国STEM毕业生短缺的问题,其中许多人最初表示对STEM领域感兴趣,但后来失去了其他专业。研究经验已被证明有助于提高对STEM的持续兴趣,但它们往往只影响到少数学生,并不一定能让所有学生公平地获得机会。本项目将针对未申报的专业,建立基于课程的宏基因组学数据分析本科研究经验(CURE),以解决这些问题。在开发改进的分类方法方面,研究将集中在(1)开发算法,利用分类信息和普遍保守的标记基因,在单个样本中对特征良好的物种和不同的物种进行分类;(2)开发泛基因组感知算法,利用同一物种在多个样本中共现进行分类。这种方法将最大限度地从单个样本中收集信息,同时避免在利用来自多个样本的数据时对样本内部和样本之间的基因组守恒进行假设。开发的方法将使用模拟宏基因组以及来自土壤的真实数据进行验证。具体来说,森林火灾后将纵向跟踪土壤群落,在那里大多数微生物被杀死,然后多样性慢慢增加到基线。除了验证分类方法的性能外,这些数据还将用于调查土壤为什么能够保持如此高水平的微生物多样性的各种假设。在该项目中设计和实施的CURE课程旨在培养在所有科学分支中变得越来越重要的“大数据”分析技能,并使学生接触到霰弹枪元基因组数据的分析。该课程将与非常受欢迎的“小地球”课程协同开展,后者目前面向美国41个州和14个国家的近万名学生。“微小地球”网络将作为数据的中心存储,可供缺乏霰弹枪测序资源的机构使用。随着时间的推移,这个数据存储库将促进对“大问题”的探索,从而使学生主导的发现成为众包。欲了解更多信息,请访问http://jason-c-kwan.github.io/CAREER_results.This该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Almost every environment on Earth is home to communities of microbes, whose metabolism profoundly influences all other life. Understanding such communities is therefore important to many fields, such as agriculture, biogeochemistry, oceanography, biology, ecology, and so on. However, the overwhelming majority of environmental microbes have never been isolated and grown in the laboratory. Therefore, a key to understanding microbial communities is the resolution of genomes from culture-independent sequencing (metagenomics), where the genomes of individual species are reconstructed from a mixed metagenome in a process called "binning". However, inaccurate binning hampers the ability to gain fundamental insights into microbial communities, for example in situations where it is uncertain whether all of a particular genome has been recovered from the dataset, or that all sequences in a bin are really from the same genome. In this regard, existing binning methods have various problems, including (1) poor performance with highly complex metagenomes, (2) assumption that all input sequences are bacterial, leading to impure bins, (3) lack of internal validation that bins make biological sense, and (4) ignoring the presence of multiple closely related strains and genome variability within a species (the "pangenome" concept). The goals of this project are to develop a binning method that can handle the novelty and complexity of even the most diverse microbial communities on Earth, such as soil. The methods developed have the potential to transform the study of microbial communities, enabling the determination of "who is doing what?" even in highly complex communities associated with higher organisms with unknown genomes. With increased understanding of microbial communities through genome-resolved metagenomics, it will eventually be possible to model and predict complex emergent behavior of communities and delineate how their capabilities are different from the component organisms in isolation. This project also aims to address the current shortfall of STEM graduates in the United States, many of whom initially express interest in STEM fields but then are lost to other majors. Research experiences have been shown to help increase persistent interest in STEM, but they often reach only a few students, and do not necessarily allow all students equitable access. This project will establish a course-based undergraduate research experience (CURE) in the analysis of metagenomics data, aimed at undeclared majors, to address these problems.In developing improved binning methods, the research will focus on (1) developing algorithms to leverage taxonomy information and universally-conserved marker genes for binning both well-characterized and divergent species in single samples, and (2) developing pangenome-aware algorithms to leverage co-occurrence of the same species in multiple samples for binning. This approach will maximize the information that will be gleaned from single samples, while avoiding assumptions on genome conservation within and between samples when leveraging data from multiple samples. The developed methods will be validated using simulated metagenomes as well as real data from soil. Specifically, soil communities will be followed longitudinally after forest fires, where most microbes are killed and then diversity slowly increases to the baseline. In addition to validating the performance of binning methods, these data will be used to investigate various hypotheses on why soil is able to maintain such high levels of microbial diversity. The CURE course which will be designed and implemented during this project will aim to develop skills in "big data" analysis that are becoming increasingly important in all branches of science and expose students to the analysis of shotgun metagenome data. The course will synergize with the highly popular Tiny Earth course, which is currently offered to nearly 10,000 students in 41 US states and 14 countries. The Tiny Earth network will act as a central deposit for data that can be used by institutions that lack resources for shotgun sequencing. Over time this data repository will facilitate the exploration of "big questions", thus crowdsourcing student-led discovery. For further information, visit http://jason-c-kwan.github.io/CAREER_results.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1128/mbio.02030-21
发表时间:
2022-04-26
期刊:
mBio
影响因子:
6.4
作者:
[]
通讯作者:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位: