Mapping short reads in RY-space: a novel strategy for extending the phylogenetic range of Next Generation Sequence mapping algorithms
Mapping short reads in RY-space: a novel strategy for extending the phylogenetic range of Next Generation Sequence mapping algorithms
批准号:
BB/I02347X/1
负责人:
Neil Hall
金额:
$15.2万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2012
资助国家:
英国
项目状态:
已结题
起止时间:
2012 至 --
中文摘要
现代DNA测序仪可以产生数十亿个通常只有50到100个分子(核苷酸)大小的短DNA片段。为了理解这些片段,通常需要将它们与先前测序的DNA进行比较,通常是整个基因组,长度可能有数十亿个核苷酸。这是一个很难计算的挑战。虽然比较DNA片段是一个相对简单的过程,但是产生的大量片段,加上参考基因组的大小,意味着如果没有精心设计的软件,用现有的计算机在实际的时间框架内处理一次测序产生的所有数据是很困难的,如果不是不可能的话。快速和高效的软件已经适当地开发出来,但是,为了在最新的计算机上达到合理的绘图速度,必须做出妥协。目前的软件只能适应短DNA序列之间的一些差异,以成功识别匹配(通常在短时间内只有两到三个差异是可能的)。正是这种限制使得比较物种之间的DNA变得非常困难,因为物种之间的差异是可以预料到的,因此,如果所讨论的生物体之前没有被测序,就很难理解短读数据。不幸的是,大多数具有研究、经济和临床重要性的物种尚未被测序。以前没有被测序的物种必须经历一个昂贵得多的长读基因组测序过程,这意味着对许多生物的基因组分析在目前的技术下在经济上是不可行的。我们提出了一种绘制DNA的新方法,至少在一定程度上克服了这一困难。我们的方法是基于观察DNA序列中存在的进化模式,当信息含量降低和序列模式简化时,更充分地揭示了DNA序列。一段DNA可以被认为是四种核苷酸的复杂模式:腺嘌呤、胞嘧啶、鸟嘌呤和胸腺嘧啶(简称A、C、G和T)。A和G是嘌呤核苷酸,而C和T是嘧啶。人们早就认识到,随着生物体的进化,一种嘌呤突变为另一种嘌呤,或一种嘧啶突变为另一种嘧啶的速度往往高于嘌呤突变为嘧啶的速度,反之亦然。这种突变率的不平衡将在DNA序列中产生嘌呤和嘧啶的模式,这些模式是物种之间共同祖先的更稳定的基序,而不是更嘈杂的核苷酸模式。我们将使用这种更强大的模式来匹配来自不同物种的序列:我们将开发软件,将DNA序列简化到嘌呤和嘧啶含量,使用与目前使用的原始DNA序列相同的方法来比较它们以确定相似性,然后将它们转换回原始核苷酸用于后续分析。因为从单个核苷酸到嘌呤/嘧啶身份的转换很简单,翻译reads的绘制速度将与原始DNA相当。因此,使用我们的策略,映射应该几乎和当前的方法一样快,并且使用相似水平的计算机资源。然而,一个物种可以与另一个物种进行比较的程度将会大得多,这意味着即使没有特定物种的参考序列,也可以用低成本的短读测序技术对更多的生物进行测序。因此,较低成本的测序将比目前可能的更广泛的生物范围变得实用,确保目前正在开发的依赖于短读段测序的新技术可以应用于更多的环境。
英文摘要
Modern DNA sequencers can generate billions of short DNA fragments that will typically be only 50 to 100 molecules (nucleotides) in size. To make sense of these fragments, or reads, it is usually necessary to compare them with previously sequenced DNA, often an entire genome, which may be several billions of nucleotides in length. This is a difficult computational challenge. Although comparing one piece of DNA with another is a relatively simple process, the sheer number of fragments generated, combined with the often large size of the reference genome, means that without well designed software it is hard, if not impossible, to process all the data produced by a single sequencing run within a practical timeframe using current computers. Fast and efficient software have been duly developed but, in order to achieve a reasonable mapping speed with the latest computers, compromises have to be made. Current software can accommodate only a few differences between short DNA sequences for a successful match to be identified (often only two or three differences within a short stretch is possible). It is this limitation which makes it very difficult to compare DNA between species, where a high level of variation is to be expected, and so it is hard to make sense of short-read data if the organism in question has not been sequenced before. Unfortunately most species of research, economic, and clinical importance have yet to be sequenced. Species that haven't been sequenced before must go through a far more expensive process of long-read genome sequencing, and this means that the genomic analysis of many organisms remain financially unfeasible with current technologies. We propose a new way of mapping DNA that will, at least in part, overcome this difficulty. Our approach is based on the observation that evolutionary patterns existing within DNA sequences are more fully revealed when the information content is reduced and the sequence pattern simplified. A stretch of DNA can be thought of as a complex pattern of four types of nucleotide: Adenine, Cytosine, Guanine and Thymine (A, C, G, and T for short). A and G are purine nucleotides, whilst C and T are pyrimidines. It has long been recognised that, as organisms evolve, the rate at which a purine mutates into another purine, or a pyrimidine to another pyrimidine, will tend to be higher than when purines mutate into pyrimidines and visa versa. This imbalance in mutation rate will create patterns of purines and pyrimidines within DNA sequences that are more stable motifs of shared ancestry between species than is the case with more noisy nucleotide patterns. We will use this more robust pattern to match sequences together from different species: we will develop software which will simplify DNA sequences down to their purine and pyrimidine content alone, compare them to identify similarities using approaches equivalent to that currently used with raw DNA sequences, then convert them back into their original nucleotides for subsequent analysis. Because the conversion from individual nucleotides to their purine/pyrimidine identities alone is simple, the speed with which translated reads can be mapped will be comparable to that achieved with raw DNA. Thus, using our strategy, mapping should almost be as quick as current methods and use similar levels of computer resources. However, the extent to which one species can be compared with another will be far greater meaning that it will be possible to sequence more organisms with low-cost short-read sequencing technologies even when a reference sequence for that particular species is not available. As a result lower cost sequencing will become practical for a much wider range of organisms than is currently possible, ensuring that the new techniques currently being developed, that rely on short read sequencing, can be applied in many more contexts.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1093/nar/gku1341
发表时间:
2015-03-31
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Schirmer M, Ijaz UZ, D'Amore R, Hall N, Sloan WT, Quince C]
通讯作者:
Quince C
Analysis of the bread wheat genome using whole-genome shotgun sequencing.
使用全基因组shot弹枪测序分析面包小麦基因组。
DOI:
10.1038/nature11650
发表时间:
2012-11-29
期刊:
Nature
影响因子:
64.8
作者:
[]
通讯作者:
Open Access Block Award 2024 - Earlham Institute
-
批准号:EP/Z531492/1
-
项目类别:Research Grant
-
资助金额:$2.37万
-
财政年份:2024
-
负责人:Neil Hall
-
依托单位:
Open Access Block Award 2023 - Earlham Institute
-
批准号:EP/Y529126/1
-
项目类别:Research Grant
-
资助金额:$1.54万
-
财政年份:2023
-
负责人:Neil Hall
-
依托单位:
Open Access Block Award 2022 - Earlham Institute
-
批准号:EP/X526095/1
-
项目类别:Research Grant
-
资助金额:$1.72万
-
财政年份:2022
-
负责人:Neil Hall
-
依托单位:
ELIXIR-UK Coordination Office
-
批准号:BB/X011100/1
-
项目类别:Research Grant
-
资助金额:$74.16万
-
财政年份:2022
-
负责人:Neil Hall
-
依托单位:
The Earlham Institute 2021 Flexible Talent Mobility Account
-
批准号:BB/W510890/1
-
项目类别:Research Grant
-
资助金额:$13.76万
-
财政年份:2021
-
负责人:Neil Hall
-
依托单位:
Business Case for a Catalyst Partnership in Artificial Intelligence between the Alan Turing Institute and the Norwich Biosciences Institutes
-
批准号:BB/V509267/1
-
项目类别:Research Grant
-
资助金额:$76.45万
-
财政年份:2020
-
负责人:Neil Hall
-
依托单位:
Development of single-cell sequencing technology for microbial populations
-
批准号:BB/R022526/1
-
项目类别:Research Grant
-
资助金额:$19.06万
-
财政年份:2018
-
负责人:Neil Hall
-
依托单位:
Ultra High-Throughput Sequencing for Norwich Research Park and the UK National Capability in Genomics
-
批准号:BB/R014329/1
-
项目类别:Research Grant
-
资助金额:$97.84万
-
财政年份:2018
-
负责人:Neil Hall
-
依托单位:
Earlham Institute UKRI Innovation Fellowships: BBSRC Flexible Talent Mobility Accounts
-
批准号:BB/R50659X/1
-
项目类别:Research Grant
-
资助金额:$11.47万
-
财政年份:2017
-
负责人:Neil Hall
-
依托单位:
Wheat Pan-Genomics
-
批准号:BB/P010768/1
-
项目类别:Research Grant
-
资助金额:$138.2万
-
财政年份:2017
-
负责人:Neil Hall
-
依托单位:
EUROPEAN PARTNERING AWARD: ELIXIR - Broadening UK Participation
-
批准号:BB/P026001/1
-
项目类别:Research Grant
-
资助金额:$2.58万
-
财政年份:2017
-
负责人:Neil Hall
-
依托单位:
ELIXIR-UK Coordination Office
-
批准号:BB/P017193/1
-
项目类别:Research Grant
-
资助金额:$97.78万
-
财政年份:2016
-
负责人:Neil Hall
-
依托单位:
Establishing a single cell genomics facility
-
批准号:BB/M012638/1
-
项目类别:Research Grant
-
资助金额:$41.78万
-
财政年份:2015
-
负责人:Neil Hall
-
依托单位:
Establishing Single Molecule Real Time Sequencing for the North of the UK
-
批准号:BB/L014777/1
-
项目类别:Research Grant
-
资助金额:$45.09万
-
财政年份:2014
-
负责人:Neil Hall
-
依托单位:
Centre for Genomic Research: Genomics Hub Renewal
-
批准号:MR/K002279/1
-
项目类别:Research Grant
-
资助金额:$64.22万
-
财政年份:2013
-
负责人:Neil Hall
-
依托单位:
Development and benchmarking of improved computational methods for transcript-level expression analysis using RNA-seq data
-
批准号:BB/J007994/1
-
项目类别:Research Grant
-
资助金额:$40.26万
-
财政年份:2012
-
负责人:Neil Hall
-
依托单位:
Large scale molecular haplotyping using next generation sequencing
-
批准号:BB/I004416/1
-
项目类别:Research Grant
-
资助金额:$54.45万
-
财政年份:2011
-
负责人:Neil Hall
-
依托单位:
Mining the allohexaploid wheat genome for useful sequence polymorphisms
-
批准号:BB/G013004/1
-
项目类别:Research Grant
-
资助金额:$135.4万
-
财政年份:2009
-
负责人:Neil Hall
-
依托单位:
High throughput Sequencing Hub for the North of England
-
批准号:G0900753/1
-
项目类别:Research Grant
-
资助金额:$288.99万
-
财政年份:2009
-
负责人:Neil Hall
-
依托单位:
国内基金
海外基金
登录
查看更多内容
ESL1(Erect and Short Leaf 1)调控谷子株型的分子机制解析
-
批准号:32301849
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2023
-
负责人:张伟
-
依托单位:
Long-TSLP和Short-TSLP佐剂对新冠重组蛋白疫苗免疫应答的影响与作用机制
-
批准号:--
-
项目类别:面上项目
-
资助金额:58万元
-
批准年份:2021
-
负责人:叶亮
-
依托单位:
与SHORT-ROOT和SCARECROW发育途径相关的IDD家族基因的确定和功能研究
-
批准号:31871493
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2018
-
负责人:Hongchang Cui
-
依托单位:
long-TSLP和short-TSLP调控肺成纤维细胞有氧糖酵解在哮喘气道重塑中的作用和机制研究
-
批准号:81700034
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:余常辉
-
依托单位:
哮喘气道上皮来源long-TSLP/short-TSLP失衡对气道重塑中成纤维细胞活化的分子机制研究
-
批准号:81670026
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2016
-
负责人:蔡绍曦
-
依托单位:
短链脂肪酸上调小肠上皮紧密连接屏障功能的机制
-
批准号:31040041
-
项目类别:专项基金项目
-
资助金额:10.0万元
-
批准年份:2010
-
负责人:王鹏远
-
依托单位:
MBR中溶解性微生物产物膜污染界面微距作用机制定量解析
-
批准号:50908133
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2009
-
负责人:梁爽
-
依托单位:
高通量DNA测序片段的拼接
-
批准号:30871393
-
项目类别:面上项目
-
资助金额:35.0万元
-
批准年份:2008
-
负责人:陆祖宏
-
依托单位:
短QT综合征新致病基因的定位研究
-
批准号:30771183
-
项目类别:面上项目
-
资助金额:8.0万元
-
批准年份:2007
-
负责人:吕利雄
-
依托单位:
基于短寿蛋白肿瘤疫苗诱导的抗瘤作用及其机制的研究
-
批准号:30771999
-
项目类别:面上项目
-
资助金额:33.0万元
-
批准年份:2007
-
负责人:王立新
-
依托单位: