课题基金 / 基金详情

COMPUTATIONAL METHODS FOR MICROBIAL NEXT GENERATION RE-SEQUENCING DATA

COMPUTATIONAL METHODS FOR MICROBIAL NEXT GENERATION RE-SEQUENCING DATA
微生物下一代重测序数据的计算方法
批准号:
BB/M001121/1
负责人:
David Robertson
金额:
$34.93万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2014
资助国家:
英国
项目状态:
已结题
起止时间:
2014 至 --

项目摘要

项目成果

David Robertson的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The overwhelming majority of life that has existed or exists is invisible to the naked eye (collectively termed the microbes, or microorganisms) and, including the viruses, forms large and complex communities. Characterising the species present, genome composition and genetic variation in these communities has been a major focus of 'metagenomics', the genomic study of mixed samples from the environment, or from animals or humans, for example, from an animal's gut or a soil microbial ecosystems. Contemporary sequencing technologies (next generation sequencing, NGS) have massively parallelized the determination of nucleotide order within genetic material resulting in our ability to rapidly sequence different microbes. This introduces the potential to explore microbial communities and genetic diversity on a scale that was previously unprecedented. Computational methods play a central role in the analysis, alignment and assembly of NGS data. However, the amount of data being generated is outstripping our ability to analyse them routinely, let alone carry out appropriate comparative analysis. This lack of software arises because most research effort is being directed at assembling single complete genomes from next generation sequence data. However, with microbes many interesting questions concern the diversity of sequences present in a community and population variation, revealed by 'ultra-deep' sequencing. Emerging approaches aim to build a de novo assembly of the NGS reads (each read is an individual sequence fragment corresponding to a region of a genome) in a similar fashion to a jigsaw puzzle where a picture is constructed by joining all the matching pieces together. In de novo assembly the genome sequence is constructed by allocating matching short reads together. The majority of the existing de novo assembly approaches for NGS data make extensive use of the de Bruijn graph method. However, building de Bruijn graphs for very large NGS data sets is very demanding because they require hefty computational resources. In this project we propose to develop novel computational methods, based on compressing the individual NGS reads by recasting them as numerical sequences (and working with this transformed/compressed data directly) that will be generically useful for all types of microbial data sets. In order to do this we will explore novel methods for representing short-read sequence data graphically and apply established mathematical approaches for efficient data mining. The particular problem we will address is the assembly of NGS data sets where the variation in the sample needs to be considered in the analysis. In metagenomics data variation between reads corresponds to both distinct microbial species and variation within individual species or viral populations. A particularly important focus is the ability to assembly a genome without a reference sequence for comparison (de novo assembly) as an appropriate reference genome is frequently not available for many microbes and, even when a reference is available, genome architecture can vary within a species.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
Marginalised stack denoising autoencoders for metagenomic data binning
用于宏基因组数据分箱的边缘化堆栈去噪自动编码器
DOI: 10.1109/cibcb.2017.8058552
发表时间: 2017
期刊:
影响因子: --
作者: [Kouchaki S]
通讯作者: Kouchaki S
DOI: 10.1109/ssci.2016.7849955
发表时间: 2016-12
期刊: 2016 IEEE Symposium Series on Computational Intelligence (SSCI)
影响因子: --
作者: [S. Kouchaki;Santosh Tirunagari;Avraam Tapinos;D. Robertson]
通讯作者: S. Kouchaki;Santosh Tirunagari;Avraam Tapinos;D. Robertson
Alignment by numbers: sequence assembly using compressed numerical representations
按数字对齐:使用压缩数字表示进行序列组装
DOI: 10.1101/011940
发表时间: 2014
期刊:
影响因子: --
作者: [Tapinos A]
通讯作者: Tapinos A
DOI: 10.1093/ve/vew022
发表时间: 2016-07
期刊: Virus evolution
影响因子: 5.3
作者: [Rose R, Constantinides B, Tapinos A, Robertson DL, Prosperi M]
通讯作者: Prosperi M
6
    Integrative viral genomics and bioinformatics platform
    • 批准号:
      MC_UU_00034/5
    • 项目类别:
      Intramural
    • 资助金额:
      $1082.69万
    • 财政年份:
      2023
    • 负责人:
      David Robertson
    • 依托单位:
    ISCF HDRUK DIH Sprint Exemplar: Graph-Based Data Federation for Healthcare Data Science
    • 批准号:
      MC_PC_18029
    • 项目类别:
      Intramural
    • 资助金额:
      $33.14万
    • 财政年份:
      2019
    • 负责人:
      David Robertson
    • 依托单位:
    Capital Award in Support of Early Career Researchers: "Edinburgh Vishub"
    • 批准号:
      EP/S018042/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $25.48万
    • 财政年份:
      2019
    • 负责人:
      David Robertson
    • 依托单位:
    eBase: Evidence-Base; growing the Big Grant Club
    • 批准号:
      EP/S012087/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $74.17万
    • 财政年份:
      2018
    • 负责人:
      David Robertson
    • 依托单位:
    国内基金
    海外基金
    Computational Methods for Analyzing Toponome Data