Computational Techniques for Advancing Untargeted Metabolomics Analysis
Computational Techniques for Advancing Untargeted Metabolomics Analysis
批准号:
10394012
负责人:
Soha Hassoun
金额:
$1.09万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-09-23 至 2023-08-31
关键词:
AddressAdoptionBiologicalBiomedical ResearchBlood CirculationCase StudyChemical StructureChemicalsComplexComputational TechniqueComputing MethodologiesConsumptionDataData SetDatabasesDevelopmentDiseaseEngineeringEnsureFeedbackGoalsHealthHumanInternetIntestinesLabelLettersLiteratureMachine LearningMapsMass Spectrum AnalysisMeSH ThesaurusMeasurementMeasuresMetabolicMetabolismMethodsModelingMolecularMolecular StructureNutritionalOrganPathway interactionsPerformancePlayProbabilityPropertyPubChemPubMedPublic DomainsResearchResearch PersonnelRoleRunningSamplingStatistical ModelsStructureSurveysTechniquesTestingTimeTissuesTrainingUncertaintyValidationWorkannotation systembasebiomarker discoverychemical standardcombinatorialcomputerized toolscostdark matterdeep learningdesigndrug developmentdrug discoveryexperimental studygastrointestinal systemgut microbiotainterestlarge datasetsmetabolomemetabolomicsmicrobiotamicrobiota metabolitesneural networknovelnutritionopen sourcephysical propertysmall moleculetool
中文摘要
项目摘要/摘要
利用质谱学(MS)检测和定量细胞代谢产物已经显示
在生物标记物发现、营养分析和其他生物医学研究领域前景广阔。尽管最近
随着分析技术的进步,我们解释MS测量结果的能力仍然有限。最大的
代谢组学中的挑战是注释,其中被测量的化合物被指定为化学身份。这个
当前计算工具的注释率较低。在几项调查的代谢组学研究中,不到
所有化合物中有20%有注释。导致注释率低的另一个因素是缺乏系统性
设计候选集的方法,即可在注释过程中使用的假定化学同一性的列表。
考虑到分子的组合空间很大,依赖现有的数据库是有问题的
有许多与生物相关的化合物没有在数据库中编目或在
文学作品。第二个但也是重要的挑战是解释测量结果以了解新陈代谢
正在研究的样本的活性。目前的技术在利用关于信息的复杂信息方面受到限制
用于阐明代谢活性的样本。
这个项目的目标是开发计算技术来推进大尺度地震的解释
代谢组学测量。为了应对当前的挑战,我们建议追求三个目标:(1)工程学
增强生物发现的候选集合。(2)开发新的标注技术,包括使用
深度学习和增量构建方法,以推荐最好地解释
测量。(3)构建分析代谢活动的概率模型。每项技术都将是
使用化学标准进行了严格的计算和实验验证。两个详细的案例研究
肠道微生物区系将使我们能够进一步验证我们的工具。微生物区系衍生的代谢物已经被
在循环中被检测到,并被证明在消化以外的器官和组织中参与宿主细胞通路
系统。因此,识别这些代谢物对于了解微生物区系的代谢功能至关重要。
并阐明它们的作用机制。复杂的测试用例将挑战我们的技术,提供反馈
在开发过程中,并允许我们进一步传播我们的技术。我们将与早期采用者密切合作
我们的工具,如支持信中所建议的,以进一步验证我们的工具并鼓励广泛采用。全
提议的工具将是开源的,并可通过网络访问。我们的工具承诺改变当前
解释新陈代谢组学数据的实践超越了目前数据库的可能,最新注释
工具、统计和过度表述分析或其组合。利用机器学习和大数据
本文提出的数据集定义了代谢组学分析中最有前途的研究方向。
英文摘要
PROJECT SUMMARY/ABSTRACT
Detecting and quantifying products of cellular metabolism using mass spectrometry (MS) has already shown
great promise in biomarker discovery, nutritional analysis and other biomedical research fields. Despite recent
advances in analysis techniques, our ability to interpret MS measurements remains limited. The biggest
challenge in metabolomics is annotation, where measured compounds are assigned chemical identities. The
annotation rates of current computational tools are low. For several surveyed metabolomics studies, less than
20% of all compounds are annotated. Another contributing factor to low annotation rates is the lack of systematic
ways of designing a candidate set, a listing of putative chemical identities that can be used during annotation.
Relying on exiting databases is problematic as considering the large combinatorial space of molecular
arrangements, there are many biologically relevant compounds not catalogued in databases or documented in
the literature. A secondary yet important challenge is interpreting the measurements to understand the metabolic
activity of the sample under study. Current techniques are limited in utilizing complex information about the
sample to elucidate metabolic activity.
The goal of this project is to develop computational techniques to advance the interpretation of large-scale
metabolomics measurements. To address current challenges, we propose to pursue three Aims: (1) Engineering
candidate sets that enhance biological discovery. (2) Developing new techniques for annotation including using
deep learning and incremental build out methods to recommend novel chemical structures that best explain the
measurements. (3) Constructing probabilistic models to analyze metabolic activity. Each technique will be
rigorously validated computationally and experimentally using chemical standards. Two detailed case studies on
the intestinal microbiota will allow us to further validate our tools. Microbiota-derived metabolites have been
detected in circulation and shown to engage host cellular pathways in organs and tissues beyond the digestive
system. Identifying these metabolites is thus critical for understanding the metabolic function of the microbiota
and elucidating their mechanisms. The complex test cases will challenge our techniques, provide feedback
during development, and allow us to further disseminate our techniques. We will work closely with early adopters
of our tools, as proposed in supporting letters, to further validate our tools and encourage wide adoption. All
proposed tools will be open source and made accessible through the web. Our tools promise to change current
practices in interpreting metabolomics data beyond what is currently possible with databases, current annotation
tools, statistical and overrepresentation analysis, or combinations thereof. The use of machine learning and large
data sets as proposed herein defines the most promising research direction in metabolomics analysis.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Using Common Fund Datasets to Illuminate Drug-Microbial Interactions
-
批准号:10777339
-
项目类别:
-
资助金额:$30.07万
-
财政年份:2023
-
负责人:Soha Hassoun
-
依托单位:
Deep Learning Models for Metabolomics Analysis
-
批准号:10552395
-
项目类别:
-
资助金额:$21.67万
-
财政年份:2023
-
负责人:Soha Hassoun
-
依托单位:
Computational Techniques for Advancing Untargeted Metabolomics Analysis
-
批准号:10022125
-
项目类别:
-
资助金额:$37.9万
-
财政年份:2019
-
负责人:Soha Hassoun
-
依托单位:
Computational Techniques for Advancing Untargeted Metabolomics Analysis
-
批准号:10242075
-
项目类别:
-
资助金额:$37.21万
-
财政年份:2019
-
负责人:Soha Hassoun
-
依托单位:
Computational Techniques for Advancing Untargeted Metabolomics Analysis
-
批准号:10480818
-
项目类别:
-
资助金额:$37.25万
-
财政年份:2019
-
负责人:Soha Hassoun
-
依托单位:
海外基金