Biology-aware machine learning methods for characterizing microbiome genotype and phenotype
Biology-aware machine learning methods for characterizing microbiome genotype and phenotype
批准号:
10798957
负责人:
Siavash Mir arabbaygi
金额:
$15.1万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-09-15 至 2026-08-31
关键词:
AdoptedAlgorithmsAreaAwardAwarenessBiologicalBiologyCharacteristicsComputing MethodologiesDataData SetEnvironmentGenotypeGrantKnowledgeLaboratoriesMachine LearningMeasurableMethodsModelingModernizationOrganismPhenotypePhylogenetic AnalysisPhylogenyProcessRecording of previous eventsResearchSamplingSequence AlignmentServicesShapesStatistical MethodsTechniquesTestingTrainingWorkdesigngenome-widegenomic dataimprovedinterestlarge datasetsmachine learning methodmicrobiomemicrobiome analysismultiple data sourcesstatistics
中文摘要
项目摘要
Mirarab实验室设计用于回答生物学和生物医学问题的计算方法,
由于可扩展性和准确性。这些方法跨越几个领域(例如,微生物组分析,多种
序列比对和基因组学),它们之间的共同点是进化建模。更
最近,许多开发的方法是基于机器学习的。该实验室开发了可扩展的,
用于重建进化历史的精确方法(即,这些历史,在当下,
流式生物医学应用。本实验室开发的方法(例如,(1),(2),(3),(4),(5),(6),(7),(9),(10),(11),(12),(13),(14),(15),(16),(17),(19)。
现代全基因组遗传学的前沿虽然该实验室以前更专注于推断物种
历史,通过MIRA赠款,它已将重点转移到开发微生物组分析方法,
给他们带来了一系列独特的挑战。
作为MIRA应用程序的一部分,Mirarab实验室将专注于设计,测试和应用改进的
微生物组数据的统计分析方法。这些方法将针对两个问题。(i)轮廓:
什么样的生物体构成了一个给定的样本?(ii)协会:样品在其有机体中有何不同
这些差异如何与其环境的可测量特征联系起来?而
这两个问题都经过了大量的研究,但仍然存在许多计算挑战,
一个更好的方法产生重大影响的机会。与其仅仅关注新算法,
该实验室还将致力于建立更好的参考数据集,并将多个来源的数据结合起来。因此
该项目旨在利用前所未有的计算能力,大型可用数据集和最新进展
在机器学习方面,以显著提高最先进的水平。该项目将不使用现成的机器
黑盒学习方法相反,它开发的方法,
(e.g.,的进化关系)转化为机器学习方法,
时尚.
在MIRA裁决的范围内,这一补充请求是购买一台计算服务器。的
服务器将使实验室能够利用当今前所未有的基因组数据水平,
机器学习方法是在比现有方法更具代表性的集合上训练的。因此,在本发明中,
额外的计算能力将不仅仅是为了使分析更快:它将使使用大型
用于训练的数据集不能以其他方式使用。
英文摘要
PROJECT SUMMARY
The Mirarab laboratory designs computational methods for answering biological and biomedical questions, fo-
cusing on scalability and accuracy. These methods span several areas (e.g., microbiome profiling, multiple
sequence alignment, and phylogenomics), and a common thread among them is evolutionary modeling. More
recently, many of the developed methods are based on machine learning. The lab has developed scalable and
accurate methods for reconstructing evolutionary histories (i.e., phylogenies) and using these histories in down-
stream biomedical applications. Methods developed by this lab (e.g., ASTRAL, SEPP, DEPP) are at the fore-
fronts of modern genome-wide phylogenetics. While the lab has previously focused more on inferring species
histories, through an MIRA grant, it has shifted its focus to developing methods for microbiome analyses, which
pose their a unique set of challenges.
As part of the MIRA application, the Mirarab lab will focus on designing, testing, and applying improved
methods for statistical analyses of microbiome data. These methods will target two questions. (i) Profiling:
What organisms constitute a given sample? (ii) Association: How are samples different in their organismal
composition, and how do these differences connect to measurable characteristics of their environment? While
both questions have been subject to considerable research, many computational challenges remain, providing
an opportunity for better methods to make a significant impact. Instead of focusing solely on new algorithms,
the lab will also work on building better reference datasets and combining data from multiple sources. Thus, the
project aims to harness the unprecedented computational power, large available datasets, and recent advances
in machine learning to improve state-of-the-art dramatically. The project will not use off-the-shelf machine
learning methods in a black-box fashion. Instead, it develops methods that incorporate biological knowledge
(e.g., of the evolutionary relationships) into machine learning methods in a principled biologically-motivated
fashion.
Within the context of the MIRA award, this supplementary request is to purchase a computing server. The
server will enable the lab to take advantage of the unprecedented level of genomic data available today to build
machine learning methods that are trained on a much more representative set than existing methods. Thus,
the extra computational power will not be just in the service of making analyses faster: it will enable using large
datasets for training that could not be otherwise used.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Biology-aware machine learning methods for characterizing microbiome genotype and phenotype
-
批准号:10696960
-
项目类别:
-
资助金额:$34.47万
-
财政年份:2021
-
负责人:Siavash Mir arabbaygi
-
依托单位:
Biology-aware machine learning methods for characterizing microbiome genotype and phenotype
-
批准号:10275055
-
项目类别:
-
资助金额:$34.47万
-
财政年份:2021
-
负责人:Siavash Mir arabbaygi
-
依托单位:
Biology-aware machine learning methods for characterizing microbiome genotype and phenotype
-
批准号:10810437
-
项目类别:
-
资助金额:$1.48万
-
财政年份:2021
-
负责人:Siavash Mir arabbaygi
-
依托单位:
海外基金