Developing Advanced Algorithms to Address Major Computational Challenges in Current Microbiome Research
Developing Advanced Algorithms to Address Major Computational Challenges in Current Microbiome Research
批准号:
9270498
负责人:
Yijun Sun
金额:
$31.12万
依托单位国家:
美国
项目类别:
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-05-15 至 2019-04-30
关键词:
AddressAlgorithmsAntibioticsAreaBig DataBioinformaticsBiologicalCommunitiesComputational algorithmComputer softwareDataData AnalysesData SetDevelopmentDiseaseEpidemiologyFloodsFoundationsHealthHumanHuman MicrobiomeHuman bodyInterdisciplinary StudyKnowledgeLogisticsMachine LearningMetagenomicsMethodsMicrobeMicrobiologyModelingOralOral MicrobiologyOutcomePeriodontal DiseasesPhysiological ProcessesPlayProbioticsResearchResourcesRibosomal RNARoleSamplingStructureSystemTaxonomyTechnologyTestingTimeWorkbasecohortcomputerized toolsdesigndynamic systemepidemiology studyexperimental studyinnovationinsightinterestmicrobialmicrobial communitymicrobiomemicrobiotamultidisciplinarynovelopen sourceoral behaviororal microbiomeresponsetumor progressionweb app
中文摘要
摘要
我们提出了一项为期三年的跨学科研究计划,以解决目前面临的两个关键问题
元基因组学社区。第一个问题涉及利用数百万个16S rRNA序列准确构建和注释OTU表,这是微生物组数据分析中最重要也是最困难的问题之一。目前,它缺乏能够处理极大数据量的计算算法
序列数据和构建生物上一致的OTU表。我们提出了一种新的方法,通过在一个分析框架内利用输入和参考序列、参考标注和数据聚类结构来同时执行OTU表的构建和标注。动态数据驱动截止值被用来识别不仅与数据聚类结构一致而且与引用注释一致的OTU。当成功实施时,我们的方法通常将满足处理目前由大规模研究产生的数亿个16S rRNA读取的计算需求。第二个问题是开发从海量序列数据中提取相关信息的新方法,从而促进该领域从描述性研究转向机械性研究。我们对微生物群落动力学分析特别感兴趣,它可以为通过静态实验设计无法获得的疾病发展提供丰富的洞察力,并为开发益生菌和抗生素策略以操纵微生物群落奠定关键基础。传统上,系统动力学是通过时间进程研究来实现的。然而,由于经济和后勤方面的限制,时间进程研究通常受到所检查的样本数量和随后的时间段的限制。随着测序技术的快速发展,成千上万的样本被大规模地收集起来。这为我们提供了一个独特的机会来开发一种新的分析策略来使用静态数据而不是时间进程数据来研究微生物群落动力学。据我们所知,这是第一次使用大量的静态数据来研究微生物群落的动态方面。成功实施后,我们的方法可以有效地克服时间进程研究的采样限制,并开辟了一条新的研究途径,无需执行资源密集型时间进程研究,即可研究疾病发展背后的微生物动力学。建议的管道将在一个大型口腔微生物组数据集上进行密集测试,该数据集包含约2,600个牙龈下样本(约330万个读数)。这些分析可以显着促进我们对口腔微生物群落动态行为的理解,这些微生物群落可能有助于牙周病的发展。据我们所知,以前还没有进行过这种规模的研究口腔微生物的工作。
社区动态。我们已经组建了一个多学科团队,涵盖机器学习、生物信息学和口腔微生物学等领域的专业知识。这项工作的预期结果将是一套
微生物界及其他领域的高实用性计算工具。
英文摘要
Abstract
We propose a three-year interdisciplinary research plan to address two key issues currently facing the
metagenomics community. The first issue concerns accurate construction and annotation of OTU tables using of millions of 16S rRNA sequences, which is one of the most important yet most difficult problems inmicrobiome data analysis. Currently, it lacks computational algorithms capable of handling extremely large
sequence data and constructing biologically consistent OTU tables. We propose a novel method that performs OTU table construction and annotation simultaneously by utilizing input and reference sequences, reference annotations, and data clustering structure within one analytical framework. Dynamic data-driven cutoffs are derived to identify OTUs that are consistent not only with data clustering structure but also with reference annotations. When successfully implemented, our method will generally address the computational needs of processing hundreds of millions of 16S rRNA reads that are currently being generated by large-scale studies. The second issue concerns developing novel methods to extract pertinent information from massive sequence data, thereby facilitating the field shifting from descriptive research to mechanistic studies. We are particularly interested in microbial community dynamics analysis, which can provide a wealth of insight into disease development unattainable through a static experiment design, and lays a critical foundation for developing probiotic and antibiotic strategies to manipulate microbial communities. Traditionally, system dynamics is approached through time-course studies. However, due to economical and logistical constraints, time-course studies are generally limited by the number of samples examined and the time period followed. With the rapid development of sequencing technology, many thousands of samples are being collected in large-scale studies. This provides us with a unique opportunity to develop a novel analytical strategy to use static data, instead of time-course data, to study microbial community dynamics. To our knowledge, this is the first time that massive static data is used to study dynamic aspects of microbial communities. When successfully implemented, our approach can effectively overcome the sampling limitation of time-course studies, and opens a new avenue of research to study microbial dynamics underlying disease development without performing a resource-intensive time-course study. The proposed pipeline will be intensively tested on a large oral microbiome dataset consisting of ~2,600 subgingival samples (~330M reads). The analysis can significantly advance our understanding of dynamic behaviors of oral microbial communities possibly contributing to the development of periodontal disease. To our knowledge, no prior work has been performed on this scale to study oral microbial
community dynamics. We have assembled a multidisciplinary team that covers expertise spanning the areas of machine learning, bioinformatics, and oral microbiology. The expected outcome of this work will be a set of
computational tools of high utility for the microbiology community and beyond.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金