Biology-aware machine learning methods for characterizing microbiome genotype and phenotype
Biology-aware machine learning methods for characterizing microbiome genotype and phenotype
批准号:
10696960
负责人:
Siavash Mir arabbaygi
金额:
$34.47万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-09-15 至 2026-08-31
关键词:
AdoptedAlgorithmsAreaAwarenessBiologicalBiologyBiomedical ResearchCharacteristicsComputing MethodologiesDataData SetDiseaseEnvironmentEpidemiologyGenomeGenomicsGenotypeGoalsHigh Performance ComputingImmunologyKnowledgeLaboratoriesMachine LearningMeasurableMeasuresMetagenomicsMethodsModernizationOrganismPhenotypePhylogenetic AnalysisPhylogenyProcessRecording of previous eventsResearchSamplingSequence AlignmentShapesStatistical Data InterpretationTechniquesTestingTreesUpdateWorkcomparativedeep learningdesigngenome-wideimprovedinterestmachine learning methodmicrobiomemicrobiome analysismultiple data sourcesstatistics
中文摘要
项目总结
1 Mirarab实验室为回答生物和生物医学问题设计了领先的计算方法-
2.注重可伸缩性和准确性。这些方法跨越几个领域(例如,微生物组编程fi林,
3多序列比对和系统基因组学),其中一个共同的主线是进化模式。
4.岭南。该实验室已经开发了可扩展且准确的方法来重建进化史(即,
5个系统发育史),并在下游生物医学应用中使用这些历史。重建系统发育是一项
6基本目标和许多生物分析的先导。本实验室开发的方法(例如,Astral)
7个是现代全基因组系统发育的前沿。此外,生物医学研究越来越多地使用
8个不同领域的进化史,如微生物组分析、免疫学、流行病学和比较
9基因组学。虽然该实验室以前更专注于推断物种历史,但它最近开始
10将重点转向开发微生物组分析方法。进化论的推论和应用
11分析环境微生物组样本的历史提出了一系列独特的挑战。
12在接下来的几年里,fi实验室将专注于设计、测试和应用改进的方法
13微生物组数据的统计分析。这些方法将针对两个问题。(I)ProfiLing:什么生物
14是否构成给定样本?(Ii)协会:样本的有机成分有何不同;及
15这些差异如何与其环境的可测量特征相联系?虽然这两个问题
16已经进行了大量的研究,许多计算挑战仍然存在,提供了一个机会
更好的方法使fi不能产生明显的影响。该实验室将不再只专注于新的算法,而是
18还致力于建立更好的参考数据集和合并来自多个来源的数据。因此,该项目
19旨在利用前所未有的计算能力、大量可用数据集和最新进展
20机器学习显著提高了最先进的水平。该项目不会使用现成的机器学习
21种方法,以黑盒方式进行。相反,它开发了结合生物学知识的方法(例如,
22进化关系)以一种原则性的生物激励的方式转化为机器学习方法。
23实验室将在支持fiLing和协会问题上追求几个雄心勃勃的目标。该项目将
24(I)创建方法以推断持续更新的参考比对和包含所有测序的树
25个原核生物基因组(目前有50万个)将用于前fi检测,(Ii)建立超灵敏SAM-LING的方法。
26 PROfiLING,(Iii)使用深度学习来连接使用扩增子测序和元基因组学获得的数据,
27(Iv)建立样本区分的不一致意识的系统发育测量,以及(V)发展机器学习
28种将PROfi与感兴趣的表型(如疾病)相关联的方法导致微生物组。这些新方法
29将利用统计学、机器学习、离散优化和高性能计算。一致
30有了米拉的目标,该项目可能会探索新的不可预见的机会,如果他们fi它的总体目标。
英文摘要
PROJECT SUMMARY
1 The Mirarab laboratory designs leading computational methods for answering biological and biomedical ques-
2 tions, focusing on scalability and accuracy. These methods span several areas (e.g., microbiome profiling,
3 multiple sequence alignment, and phylogenomics), and a common thread among them is evolutionary mod-
4 eling. The lab has developed scalable and accurate methods for reconstructing evolutionary histories (i.e.,
5 phylogenies) and using these histories in downstream biomedical applications. Reconstructing phylogenies is a
6 fundamental goal and a precursor to many biological analyses. Methods developed by this lab (e.g., ASTRAL)
7 are at the forefronts of modern genome-wide phylogenetics. Moreover, biomedical research increasingly uses
8 evolutionary histories in diverse areas like microbiome analyses, immunology, epidemiology, and comparative
9 genomics. While the lab has previously focused more on inferring species histories, it has recently started
10 to shift its focus to developing methods for microbiome analyses. The inference and the use of evolutionary
11 histories in analyzing environmental microbiome samples present a unique set of challenges.
12 In the next five years, the Mirarab lab will focus on designing, testing, and applying improved methods for
13 statistical analyses of microbiome data. These methods will target two questions. (i) Profiling: What organisms
14 constitute a given sample? (ii) Association: How are samples different in their organismal composition, and
15 how do these differences connect to measurable characteristics of their environment? While both questions
16 have been subject to considerable research, many computational challenges remain, providing an opportunity
17 for better methods to make a significant impact. Instead of focusing solely on new algorithms, the lab will
18 also work on building better reference datasets and combining data from multiple sources. Thus, the project
19 aims to harness the unprecedented computational power, large available datasets, and recent advances in
20 machine learning to improve state-of-the-art dramatically. The project will not use off-the-shelf machine learning
21 methods in a black-box fashion. Instead, it develops methods that incorporate biological knowledge (e.g., of the
22 evolutionary relationships) into machine learning methods in a principled biologically-motivated fashion.
23 The lab will pursue several ambitious goals for both profiling and association questions. The project will
24 (i) create methods to infer a continuously-updated reference alignment and tree encompassing all sequenced
25 prokaryotic genomes (half a million currently) to be used for profiling, (ii) build methods for ultra-sensitive sam-
26 ple profiling, (iii) use deep learning to connect data obtained using amplicon sequencing and metagenomics,
27 (iv) build discordance-aware phylogenetic measures of sample differentiation, and (v) develop machine learning
28 methods for associating a profiled microbiome to phenotypes of interest such as disease. These new methods
29 will draw on statistics, machine learning, discrete optimization, and high-performance computing. Consistent
30 with the goals of MIRA, the project may explore new unforeseen opportunities if they fit its general goals.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Biology-aware machine learning methods for characterizing microbiome genotype and phenotype
-
批准号:10275055
-
项目类别:
-
资助金额:$34.47万
-
财政年份:2021
-
负责人:Siavash Mir arabbaygi
-
依托单位:
Biology-aware machine learning methods for characterizing microbiome genotype and phenotype
-
批准号:10810437
-
项目类别:
-
资助金额:$1.48万
-
财政年份:2021
-
负责人:Siavash Mir arabbaygi
-
依托单位:
Biology-aware machine learning methods for characterizing microbiome genotype and phenotype
-
批准号:10798957
-
项目类别:
-
资助金额:$15.1万
-
财政年份:2021
-
负责人:Siavash Mir arabbaygi
-
依托单位:
海外基金