EMERALD - Enriching MEtagenomics Results using Artificial intelligence and Literature Data
EMERALD - Enriching MEtagenomics Results using Artificial intelligence and Literature Data
批准号:
BB/S009043/1
负责人:
Robert Finn
金额:
$77.25万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --
中文摘要
细菌和真菌等微生物生活在不同的环境中,包括土壤、水和人体部位,如口腔、皮肤和肠道。它们在自然界中无处不在,也表现出对极端环境的适应,如酸性矿山排水或热液喷口。长期以来,我们一直很欣赏微生物的潜力--它们对食品和饮料制造(如奶酪和啤酒)很重要,也是生物修复的关键参与者,墨西哥湾深水地平线漏油事件后,它们在分解复杂石油方面发挥的关键作用就证明了这一点。中基因组学领域提供了一个令人兴奋的机会来研究这些微生物群落,并深入了解它们存在的各个方面,即它们与人类和植物的相互作用,它们作为疾病宿主的潜力,以及作为具有生物修复或塑料回收能力的新型酶的来源。中基因组学通过直接采样环境,提取并测序它们的遗传物质(DNA),并应用计算方法来阐明微生物的组成和功能,从而研究微生物群落。这种抽样方法有助于在实验室中确定不可培养或尚未培养的微生物的特征。元基因组学实验数据通常很大(每次测序运行10-100 GB;每个项目运行100-100 GB)、复杂(包括100-1000个不同微生物)和可变的,这是由于基础实验的性质和动态种群的(子)采样。尽管关于微生物群落内的通量的知识(例如,一年中的时间或一天中的时间),元基因组数据集通常包含与样品来源或用于获得DNA和处理序列数据的方法相关的糟糕的描述(称为元数据)。为了帮助解释实验中的数据并得出有意义的生物学结论,关键是要知道两个元基因组数据集之间的差异是由于潜在的实验技术的不同还是样本的生物学性质的差异。元数据的缺乏阻碍了我们尝试应用机器学习(ML)技术来解释新的传入数据,从而阻碍了我们寻找新的生物应用的能力。为了绕过这些问题,我们的提议旨在使用不同的ML方法来丰富当前可用的元数据,并开始阐明嵌入在序列数据中的新知识。文本挖掘方法将侧重于识别关于元基因组学实验的研究文章,以挖掘和提取详细描述,这些描述将用于丰富与相应DNA序列相关联的元数据,并生成新的或改进的分类系统。这本描述词词典还将作为开发方法的模板,以发现以前未确定的元基因组学论文。我们将对这些丰富的元数据进行算法训练,以逐步了解什么标准可能应用于描述不充分的传入数据,以便在比较类似样本时确定样本来源、处理以及破译哪些实验偏差影响结果。ML方法也将用于发现新的生物功能。细菌编码基因盒,负责生产具有药物和农业价值的化合物。对构成这些录音带的基因的功能描述还不完整,而许多录音带仍有待发现。通过将ML和文本挖掘方法相结合,我们打算更好地描述这些盒式磁带,并专注于新群体的检测。支持这项工作的数据将来自关键的EMBL-EBI数据库,即EBI Metagenology和Europe PMC,以及其他资源(如MIBiG)。这里的发展将有助于解决实验数据背后的复杂性,丰富这一过程中的元数据,并为新一代可靠的预测模型奠定基础。
英文摘要
Microbes like bacteria and fungi inhabit diverse environments, including soil, water, and human body sites, such as the mouth, skin and intestine. Ubiquitous in nature, they also show adaptation to extreme environments, such as acid mine drainage or hydrothermal vents. We have appreciated the potential of microbes for a long time - they are important for food and beverage manufacturing (e.g. cheese and beer), and are key players in bioremediation, as demonstrated by their pivotal role in breaking down complex oils following the Deep Horizon oil spill in the Gulf of Mexico. The field of metagenomics offers an exciting opportunity to examine these microbial communities and gain insights into various aspects of their existence, i.e. their interaction with humans and plants, their potential as disease reservoirs, and as sources of novel enzymes with bioremediation or plastic recycling abilities.Metagenomics studies microbial communities by sampling the environments directly, extracting and sequencing their genetic material (DNA), and applying computational methods to elucidate microbial composition and function. This sampling approach helps to characterise unculturable or as yet uncultured microbes in the laboratory. Metagenomics experimental data are typically large (10-100s of GBs per sequencing run; 100s of runs per project), complex (comprising 100-1000s of different microbes) and variable due to the nature of the underlying experiments and (sub-)sampling of the dynamic populations.Despite knowledge about fluxes within a microbial community (e.g. time of year or day), metagenomic datasets typically contain poor descriptions (termed metadata) relating to the sample origin or methods used to obtain the DNA and process the sequence data. To help interpret data across experiments and derive meaningful biological conclusions, it is crucial to know whether a difference between two metagenomics datasets is due to differences in underlying experimental techniques or the biological qualities of the sample. The lack of metadata has impeded our attempts to apply machine learning (ML) techniques to interpret new incoming data, and therefore our capacity to find novel biological applications.To circumvent these issues, our proposal aims to employ different ML methodologies to enrich the currently available metadata and start elucidating new knowledge embedded in the sequence data. The text mining approach will focus on identifying research articles on metagenomics experiments to unearth and extract detailed descriptions which will be used to enrich the metadata associated with the corresponding DNA sequences and generate new or improved classification systems. This dictionary of descriptor terms will also serve as the template for developing methods to discover previously unidentified metagenomics papers. We will train algorithms on this enriched metadata to progressively learn what criteria might be applied to incoming data with inadequate descriptions in order to determine sample origin, processing, as well as decipher which experimental biases affect the results, when comparing similar samples.ML approaches will also be used for the discovery of new biological functions. Bacteria encode gene cassettes that are responsible for producing compounds of pharmaceutical and agricultural value. Functional descriptions for the genes constituting these cassettes are incomplete, while many cassettes still await discovery. By combining the ML and text mining approaches, we intend to better describe these cassettes and also focus on the detection of novel groups.Data underpinning this work will originate from key EMBL-EBI databases, namely EBI Metagenomics and Europe PMC, as well as other resources (e.g. MIBiG). Developments aimed at herein will help resolve complexities underlying experimental data, enriching the metadata in the process and also laying the foundation for a new generation of reliable predictive models.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1093/gigascience/giac077
发表时间:
2022-08-11
期刊:
GigaScience
影响因子:
9.2
作者:
[]
通讯作者:
A machine learning framework for discovery and enrichment of metagenomics metadata from open access publications
用于从开放获取出版物中发现和丰富宏基因组元数据的机器学习框架
DOI:
10.21203/rs.3.rs-1396476/v1
发表时间:
2022
期刊:
影响因子:
--
作者:
[Nassar M]
通讯作者:
Nassar M
DOI:
10.1101/2023.05.23.540769
发表时间:
2023-10
期刊:
bioRxiv
影响因子:
--
作者:
[Santiago Sanchez;Joel D. Rogers;Alexander B Rogers;Maaly Nassar;J. Mcentyre;M. Welch;F. Hollfelder;R. Finn]
通讯作者:
Santiago Sanchez;Joel D. Rogers;Alexander B Rogers;Maaly Nassar;J. Mcentyre;M. Welch;F. Hollfelder;R. Finn
Enriching MGnify Genomes to capture the full spectrum of the microbiota and bolster taxonomic classifications
-
批准号:BB/V01868X/1
-
项目类别:Research Grant
-
资助金额:$118.4万
-
财政年份:2022
-
负责人:Robert Finn
-
依托单位:
2020BBSRC-NSF/BIO: REDEFINE - Development of efficient, large-scale metagenomics sequence comparison algorithms to facilitate novel genomic insights
-
批准号:BB/W002965/1
-
项目类别:Research Grant
-
资助金额:$63.57万
-
财政年份:2022
-
负责人:Robert Finn
-
依托单位:
SENSE - Screening of ENvironmental SEquences to discover novel protein functions using informatics target selection and high-throughput validation
-
批准号:BB/T000902/1
-
项目类别:Research Grant
-
资助金额:$27.61万
-
财政年份:2020
-
负责人:Robert Finn
-
依托单位:
EBI Metagenomics - enabling the reconstruction of microbial populations
-
批准号:BB/R015228/1
-
项目类别:Research Grant
-
资助金额:$113.62万
-
财政年份:2018
-
负责人:Robert Finn
-
依托单位:
Bilateral NSF/BIO-BBSRC:A Metagenomics Exchange - enriching analysis by synergistic harmonisation of MG-RAST and the EBI Metagenomics Portal
-
批准号:BB/N018354/1
-
项目类别:Research Grant
-
资助金额:$104.13万
-
财政年份:2017
-
负责人:Robert Finn
-
依托单位:
Expanding Genome3D and disseminating the structural annotations via InterPro and PDBe
-
批准号:BB/N019172/1
-
项目类别:Research Grant
-
资助金额:$39.2万
-
财政年份:2016
-
负责人:Robert Finn
-
依托单位:
14 NSFBIO:Towards detailed and consistent function prediction from protein family databases
-
批准号:BB/N00521X/1
-
项目类别:Research Grant
-
资助金额:$58.0万
-
财政年份:2015
-
负责人:Robert Finn
-
依托单位:
EBI Metagenomics Portal - Towards a better understanding of community metabolism
-
批准号:BB/M011755/1
-
项目类别:Research Grant
-
资助金额:$82.32万
-
财政年份:2015
-
负责人:Robert Finn
-
依托单位:
Collaborative Research: Capillary Interfaces
-
批准号:0103954
-
项目类别:Standard Grant
-
资助金额:$11.9万
-
财政年份:2001
-
负责人:Robert Finn
-
依托单位:
Proposal for Exploratory Research
-
批准号:9729817
-
项目类别:Standard Grant
-
资助金额:$7.0万
-
财政年份:1997
-
负责人:Robert Finn
-
依托单位:
Mathematical Sciences: Classical and Applied Analysis
-
批准号:9400778
-
项目类别:Continuing Grant
-
资助金额:$2.0万
-
财政年份:1994
-
负责人:Robert Finn
-
依托单位:
Mathematical Sciences: Classical and Applied Analysis
-
批准号:9106968
-
项目类别:Continuing Grant
-
资助金额:$12.3万
-
财政年份:1991
-
负责人:Robert Finn
-
依托单位:
Mathematical Sciences: Classical and Applied Analysis
-
批准号:8902831
-
项目类别:Continuing Grant
-
资助金额:$7.45万
-
财政年份:1989
-
负责人:Robert Finn
-
依托单位:
Mathematical Sciences: Classical and Applied Analysis
-
批准号:8603631
-
项目类别:Continuing Grant
-
资助金额:$16.82万
-
财政年份:1986
-
负责人:Robert Finn
-
依托单位:
Mathematical Sciences: International Conference on Variational Methods for Free Surface Interfaces
-
批准号:8416414
-
项目类别:Standard Grant
-
资助金额:$3.0万
-
财政年份:1985
-
负责人:Robert Finn
-
依托单位:
Mathematical Sciences: Nonlinear Partial Differential Equations and Capillarity Theory
-
批准号:8307826
-
项目类别:Continuing Grant
-
资助金额:$11.96万
-
财政年份:1983
-
负责人:Robert Finn
-
依托单位:
Travel to Attend: 1st European Congress on Biotechnology; Interlaken, Switzerland; September 25 - 29,1978
-
批准号:7821903
-
项目类别:Standard Grant
-
资助金额:$0.09万
-
财政年份:1978
-
负责人:Robert Finn
-
依托单位:
海外基金