课题基金 / 基金详情

Mapping short reads in RY-space: a novel strategy for extending the phylogenetic range of Next Generation Sequence mapping algorithms

Mapping short reads in RY-space: a novel strategy for extending the phylogenetic range of Next Generation Sequence mapping algorithms
在 RY 空间中映射短读:扩展下一代序列映射算法系统发育范围的新策略
批准号:
BB/I02347X/1
负责人:
Neil Hall
金额:
$15.2万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2012
资助国家:
英国
项目状态:
已结题
起止时间:
2012 至 --

项目摘要

项目成果

Neil Hall的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Modern DNA sequencers can generate billions of short DNA fragments that will typically be only 50 to 100 molecules (nucleotides) in size. To make sense of these fragments, or reads, it is usually necessary to compare them with previously sequenced DNA, often an entire genome, which may be several billions of nucleotides in length. This is a difficult computational challenge. Although comparing one piece of DNA with another is a relatively simple process, the sheer number of fragments generated, combined with the often large size of the reference genome, means that without well designed software it is hard, if not impossible, to process all the data produced by a single sequencing run within a practical timeframe using current computers. Fast and efficient software have been duly developed but, in order to achieve a reasonable mapping speed with the latest computers, compromises have to be made. Current software can accommodate only a few differences between short DNA sequences for a successful match to be identified (often only two or three differences within a short stretch is possible). It is this limitation which makes it very difficult to compare DNA between species, where a high level of variation is to be expected, and so it is hard to make sense of short-read data if the organism in question has not been sequenced before. Unfortunately most species of research, economic, and clinical importance have yet to be sequenced. Species that haven't been sequenced before must go through a far more expensive process of long-read genome sequencing, and this means that the genomic analysis of many organisms remain financially unfeasible with current technologies. We propose a new way of mapping DNA that will, at least in part, overcome this difficulty. Our approach is based on the observation that evolutionary patterns existing within DNA sequences are more fully revealed when the information content is reduced and the sequence pattern simplified. A stretch of DNA can be thought of as a complex pattern of four types of nucleotide: Adenine, Cytosine, Guanine and Thymine (A, C, G, and T for short). A and G are purine nucleotides, whilst C and T are pyrimidines. It has long been recognised that, as organisms evolve, the rate at which a purine mutates into another purine, or a pyrimidine to another pyrimidine, will tend to be higher than when purines mutate into pyrimidines and visa versa. This imbalance in mutation rate will create patterns of purines and pyrimidines within DNA sequences that are more stable motifs of shared ancestry between species than is the case with more noisy nucleotide patterns. We will use this more robust pattern to match sequences together from different species: we will develop software which will simplify DNA sequences down to their purine and pyrimidine content alone, compare them to identify similarities using approaches equivalent to that currently used with raw DNA sequences, then convert them back into their original nucleotides for subsequent analysis. Because the conversion from individual nucleotides to their purine/pyrimidine identities alone is simple, the speed with which translated reads can be mapped will be comparable to that achieved with raw DNA. Thus, using our strategy, mapping should almost be as quick as current methods and use similar levels of computer resources. However, the extent to which one species can be compared with another will be far greater meaning that it will be possible to sequence more organisms with low-cost short-read sequencing technologies even when a reference sequence for that particular species is not available. As a result lower cost sequencing will become practical for a much wider range of organisms than is currently possible, ensuring that the new techniques currently being developed, that rely on short read sequencing, can be applied in many more contexts.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1093/nar/gku1341
发表时间: 2015-03-31
期刊: Nucleic acids research
影响因子: 14.9
作者: [Schirmer M, Ijaz UZ, D'Amore R, Hall N, Sloan WT, Quince C]
通讯作者: Quince C
Analysis of the bread wheat genome using whole-genome shotgun sequencing.
使用全基因组shot弹枪测序分析面包小麦基因组。
DOI: 10.1038/nature11650
发表时间: 2012-11-29
期刊: Nature
影响因子: 64.8
作者: []
通讯作者:
Open Access Block Award 2024 - Earlham Institute
  • 批准号:
    EP/Z531492/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $2.37万
  • 财政年份:
    2024
  • 负责人:
    Neil Hall
  • 依托单位:
Open Access Block Award 2023 - Earlham Institute
  • 批准号:
    EP/Y529126/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $1.54万
  • 财政年份:
    2023
  • 负责人:
    Neil Hall
  • 依托单位:
Open Access Block Award 2022 - Earlham Institute
  • 批准号:
    EP/X526095/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $1.72万
  • 财政年份:
    2022
  • 负责人:
    Neil Hall
  • 依托单位:
ELIXIR-UK Coordination Office
  • 批准号:
    BB/X011100/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $74.16万
  • 财政年份:
    2022
  • 负责人:
    Neil Hall
  • 依托单位:
国内基金
海外基金
ESL1(Erect and Short Leaf 1)调控谷子株型的分子机制解析
Long-TSLP和Short-TSLP佐剂对新冠重组蛋白疫苗免疫应答的影响与作用机制
  • 批准号:
    --
  • 项目类别:
    面上项目
  • 资助金额:
    58万元
  • 批准年份:
    2021
  • 负责人:
    叶亮
  • 依托单位:
与SHORT-ROOT和SCARECROW发育途径相关的IDD家族基因的确定和功能研究
  • 批准号:
    31871493
  • 项目类别:
    面上项目
  • 资助金额:
    60.0万元
  • 批准年份:
    2018
  • 负责人:
    Hongchang Cui
  • 依托单位:
long-TSLP和short-TSLP调控肺成纤维细胞有氧糖酵解在哮喘气道重塑中的作用和机制研究
  • 批准号:
    81700034
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    20.0万元
  • 批准年份:
    2017
  • 负责人:
    余常辉
  • 依托单位: