课题基金 / 基金详情

III: AF: Medium: Collaborative Research: Scalable and Highly Accurate Methods for Metagenomics

III: AF: Medium: Collaborative Research: Scalable and Highly Accurate Methods for Metagenomics
III:AF:中:协作研究:可扩展且高度准确的宏基因组学方法
批准号:
1513615
负责人:
Mihai Pop
金额:
$37.33万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2019-08-31

项目摘要

项目成果

Mihai Pop的其他基金

相似基金

相关文献

中文摘要
翻译
微生物群落的宏基因组研究可以产生数百万到数十亿的测序读数。在许多分析中,对这些序列进行准确的分类标记是一个关键组成部分,但由于在环境或宿主相关群落中发现的大多数生物体不容易在实验室中培养,这一事实使其复杂化。即使在可以培养的生物中,相对较少的生物被测序,甚至是部分测序。因此,许多常见的生物体在现有的已知基因组和基因数据库中基本上是缺失的。因此,为宏基因组序列提供分类标签需要将序列数据库中包含的知识外推到以前未见过的DNA字符串。简单的基于相似性的方法(例如,选择最好的数据库作为分类标签的最佳猜测)已被证明不够准确,从而导致开发更复杂的方法。进一步的发展是必要的,以处理新兴测序技术的特点,如高错误率与大量的插入和删除。迄今为止,人们对宏基因组分类群鉴定方法进行了评估,以评估它们在宏基因组样本中估计细菌分类群(种、属、科等)分布的能力。然而,不同的科学和临床环境可能需要特定类型的分析,而这种类型的评估可能不是最适合所有环境的。例如,在临床环境中,最重要的问题可能是检测是否存在特定的病原体,而在科学环境中,最有趣的问题可能是能够确定观察到的读数是否来自以前从未见过的物种。必须开发新的评估策略,专门针对应用程序领域的特定需求。该项目开发的所有方法都将被制成开源软件,免费提供给科学界的公众。研究人员每年将为来自全国各地的学生和博士后提供培训活动,并为少数民族服务机构和妇女提供外展计划。年代的大学。马里兰大学帕克分校也将提供夏季REU项目。该团队将开发一个新的框架,用于将生物用例的正式定义与评估数据集和度量标准集成在一起,以确保开发的软件能够充分满足最终用户的需求。其次,他们将开发基于标记的分类群鉴定和丰度分析的新方法,这些方法可以利用多个信息来源(例如,多个标记),并处理第三代测序技术的高错误率。这些方法将以开发TIPP的经验为基础。TIPP是该团队最近发布的一种分类分析软件包,其性能优于领先的宏基因组分类分析软件,特别是对于新序列或更长、高错误的序列。最后,他们计划开发这些方法的高性能计算实现,以便能够快速分析样品。在临床环境中,分析的速度尤其重要,因为医学治疗可能取决于该方法返回分析的速度。速度在非医疗应用中也很重要,其中更快的分析使研究人员能够对微生物群落进行更深入或更广泛的分析。
英文摘要
Metagenomic studies of microbial communities can generate millions to billions of sequencing reads. The assignment of accurate taxonomic labels to these sequences is a critical component in many analyses, but is complicated by the fact that the majority of the organisms found in environmental or host-associated communities cannot be easily cultured in a laboratory. Even among the organisms that can be cultured, relatively few have been sequenced, even partially. Thus, many commonly encountered organisms are largely absent from existing databases of known genomes and genes. Providing taxonomic labels to metagenomic sequences, thus, requires extrapolating the knowledge contained in sequence databases to previously unseen DNA strings. Simple similarity-based approaches (e.g., picking the best database hit as the best guess at the taxonomic label) have been shown to be insufficiently accurate, leading to the development of more sophisticated methods. Further developments are necessary to handle the characteristics of emerging sequencing technologies, such as high error rates with large numbers of insertions and deletions. To date, metagenomic taxon identification methods have been evaluated with respect to their ability to estimate the distribution of bacterial taxa (species, genera, families, etc.) within a metagenomic sample. Yet, different scientific and clinical settings may require specific types of analyses, and this one type of evaluation may not be the most appropriate for all settings. For example, in a clinical setting the most important question may be to detect whether a specific pathogen is present, while in a scientific setting the most interesting question may be to be able to determine if an observed read comes from a never-been-seen-before species. New evaluation strategies must be developed that specifically target the specific needs of the application domain. All the methods developed in the project will be made into open-source software that is freely available to the scientific public. Researchers will provide training activities each year with funds available to students and postdocs from around the country, and an outreach program to minority serving institutions and women?s colleges. A summer REU program will also be provided at the University of Maryland, College Park.The team will develop a new framework for integrating the formal definition of biological use-cases with evaluation datasets and metrics in order to ensure the software being developed adequately addresses the needs of the end-users. Second, they will develop new approaches for marker-based taxon identification and abundance profiling that can leverage multiple sources of information (e.g., multiple markers) as well as handle the high error rates of third-generation sequencing technologies. These approaches will build upon experience developing TIPP - a taxonomic profiling package recently published by the team that outperforms the leading metagenomic taxonomic profiling software, in particular for novel sequences, or for longer, high-error sequences. Finally they plan to develop high-performance computing implementations of these methods in order to enable rapid analysis of sample. Speed of analysis is particularly important in clinical settings where medical treatments may depend on the rate at which the method can return an analysis. Speed is also important in non-medical applications where faster analyses enable researchers to perform deeper or broader analyses of microbial communities.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
REU Site: Undergraduate Bioinformatics Research in Data Science for Genomics
III: Small: Genome Assembly Using Sparse Sequence Information
Algorithms for the Analysis of Data from Massively-parallel Genome Sequencing
III-CXT-Small: Graphs to Diversity: extracting genomic variation from sequence graphs
国内基金
海外基金
基于前瞻性队列的双酚AF联合果糖加重代谢损伤的靶向代谢组学研究
  • 批准号:
    2025JJ30049
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    王穆
  • 依托单位:
U2AF2-circMMP1信号轴促进结直肠癌进展的分子机制研究
U2AF2精氯酸甲基化调控RNA转录合成在MTAP缺失骨肉瘤T细胞耗竭中的机制研究
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    穆浩然
  • 依托单位:
BDA-366通过MYD88/NF-κB/PGC1β通路杀伤 KMT2A/AF9 AML细胞的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    15.0万元
  • 批准年份:
    2024
  • 负责人:
    吴利新
  • 依托单位: