课题基金 / 基金详情

EMERALD - Enriching MEtagenomics Results using Artificial intelligence and Literature Data

EMERALD - Enriching MEtagenomics Results using Artificial intelligence and Literature Data
EMERALD - 使用人工智能和文献数据丰富宏基因组学结果
批准号:
BB/S009043/1
负责人:
Robert Finn
金额:
$77.25万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --

项目摘要

项目成果

Robert Finn的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Microbes like bacteria and fungi inhabit diverse environments, including soil, water, and human body sites, such as the mouth, skin and intestine. Ubiquitous in nature, they also show adaptation to extreme environments, such as acid mine drainage or hydrothermal vents. We have appreciated the potential of microbes for a long time - they are important for food and beverage manufacturing (e.g. cheese and beer), and are key players in bioremediation, as demonstrated by their pivotal role in breaking down complex oils following the Deep Horizon oil spill in the Gulf of Mexico. The field of metagenomics offers an exciting opportunity to examine these microbial communities and gain insights into various aspects of their existence, i.e. their interaction with humans and plants, their potential as disease reservoirs, and as sources of novel enzymes with bioremediation or plastic recycling abilities.Metagenomics studies microbial communities by sampling the environments directly, extracting and sequencing their genetic material (DNA), and applying computational methods to elucidate microbial composition and function. This sampling approach helps to characterise unculturable or as yet uncultured microbes in the laboratory. Metagenomics experimental data are typically large (10-100s of GBs per sequencing run; 100s of runs per project), complex (comprising 100-1000s of different microbes) and variable due to the nature of the underlying experiments and (sub-)sampling of the dynamic populations.Despite knowledge about fluxes within a microbial community (e.g. time of year or day), metagenomic datasets typically contain poor descriptions (termed metadata) relating to the sample origin or methods used to obtain the DNA and process the sequence data. To help interpret data across experiments and derive meaningful biological conclusions, it is crucial to know whether a difference between two metagenomics datasets is due to differences in underlying experimental techniques or the biological qualities of the sample. The lack of metadata has impeded our attempts to apply machine learning (ML) techniques to interpret new incoming data, and therefore our capacity to find novel biological applications.To circumvent these issues, our proposal aims to employ different ML methodologies to enrich the currently available metadata and start elucidating new knowledge embedded in the sequence data. The text mining approach will focus on identifying research articles on metagenomics experiments to unearth and extract detailed descriptions which will be used to enrich the metadata associated with the corresponding DNA sequences and generate new or improved classification systems. This dictionary of descriptor terms will also serve as the template for developing methods to discover previously unidentified metagenomics papers. We will train algorithms on this enriched metadata to progressively learn what criteria might be applied to incoming data with inadequate descriptions in order to determine sample origin, processing, as well as decipher which experimental biases affect the results, when comparing similar samples.ML approaches will also be used for the discovery of new biological functions. Bacteria encode gene cassettes that are responsible for producing compounds of pharmaceutical and agricultural value. Functional descriptions for the genes constituting these cassettes are incomplete, while many cassettes still await discovery. By combining the ML and text mining approaches, we intend to better describe these cassettes and also focus on the detection of novel groups.Data underpinning this work will originate from key EMBL-EBI databases, namely EBI Metagenomics and Europe PMC, as well as other resources (e.g. MIBiG). Developments aimed at herein will help resolve complexities underlying experimental data, enriching the metadata in the process and also laying the foundation for a new generation of reliable predictive models.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1093/gigascience/giac077
发表时间: 2022-08-11
期刊: GigaScience
影响因子: 9.2
作者: []
通讯作者:
A machine learning framework for discovery and enrichment of metagenomics metadata from open access publications
用于从开放获取出版物中发现和丰富宏基因组元数据的机器学习框架
DOI: 10.21203/rs.3.rs-1396476/v1
发表时间: 2022
期刊:
影响因子: --
作者: [Nassar M]
通讯作者: Nassar M
DOI: 10.1101/2023.05.23.540769
发表时间: 2023-10
期刊: bioRxiv
影响因子: --
作者: [Santiago Sanchez;Joel D. Rogers;Alexander B Rogers;Maaly Nassar;J. Mcentyre;M. Welch;F. Hollfelder;R. Finn]
通讯作者: Santiago Sanchez;Joel D. Rogers;Alexander B Rogers;Maaly Nassar;J. Mcentyre;M. Welch;F. Hollfelder;R. Finn
Enriching MGnify Genomes to capture the full spectrum of the microbiota and bolster taxonomic classifications
  • 批准号:
    BB/V01868X/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $118.4万
  • 财政年份:
    2022
  • 负责人:
    Robert Finn
  • 依托单位:
2020BBSRC-NSF/BIO: REDEFINE - Development of efficient, large-scale metagenomics sequence comparison algorithms to facilitate novel genomic insights
  • 批准号:
    BB/W002965/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $63.57万
  • 财政年份:
    2022
  • 负责人:
    Robert Finn
  • 依托单位:
SENSE - Screening of ENvironmental SEquences to discover novel protein functions using informatics target selection and high-throughput validation
  • 批准号:
    BB/T000902/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $27.61万
  • 财政年份:
    2020
  • 负责人:
    Robert Finn
  • 依托单位:
EBI Metagenomics - enabling the reconstruction of microbial populations
  • 批准号:
    BB/R015228/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $113.62万
  • 财政年份:
    2018
  • 负责人:
    Robert Finn
  • 依托单位:
海外基金