课题基金 / 基金详情

Development and benchmarking of improved computational methods for transcript-level expression analysis using RNA-seq data

Development and benchmarking of improved computational methods for transcript-level expression analysis using RNA-seq data
使用 RNA-seq 数据进行转录水平表达分析的改进计算方法的开发和基准测试
批准号:
BB/J009415/1
负责人:
Magnus Rattray
金额:
$39.81万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2012
资助国家:
英国
项目状态:
已结题
起止时间:
2012 至 --

项目摘要

项目成果

Magnus Rattray的其他基金

相似基金

相关文献

中文摘要
翻译
在人类基因组测序完成后,科学家们惊讶地发现,编码蛋白质的基因比之前预测的要少得多。像人类这样复杂的有机体可以由相对较少的基因构成,其中一个原因是每个基因编码不止一种蛋白质。信使RNA(信使RNA)是一种中间分子,它将信息从细胞核中的基因组传递到产生蛋白质的核糖体。这些信使核糖核酸分子也被称为转录本,它们的完整互补被称为转录组。在它们成熟之前,这些转录本被编辑以形成不同蛋白质的模板。这种编辑过程被称为剪接,产生的不同转录本被称为剪接变体或异构体。转录组的额外复杂性是因为每个基因都有多个拷贝(例如在人类中有2个,在小麦中有6个),这些不同的拷贝被称为等位基因,在不同的条件下或在不同的组织中可以有不同的表达。转录组是转录本的集合,其中包括在细胞中表达的所有等位基因特异性基因亚型以及其他非编码RNA分子。剪接和等位基因的使用是基因功能可以以组织特有的方式进行调节的基本方式。因此,开发准确测量转录本表达的技术是理解细胞和组织并对其建模的必要步骤。最近开发的一项名为rna-seq的实验技术使人们能够史无前例地获得有关转录组的数据。需要计算方法来解释这些数据,这些数据以包含数百万个短RNA序列片段的列表的形式存在。这些片段很难解释,因为,例如,相同的片段可能来自大量不同的基因亚型。问题是,是哪一个?计算方法可以用来回答这个问题,并在给定这些数据的情况下推断样本中不同基因异构体的浓度。在这个项目中,我们将开发一种新的计算方法,在公开可用的自由软件中实现,它使用先进的统计程序来解决这个问题。该方法的一个重要特点是能够将推断的浓度与一定程度的不确定度联系起来,这种不确定度捕捉到了错误的技术和生物学来源,以及由于难以将片段分配给基因异构体而导致的问题的内在困难。我们将创建基准数据,使我们能够评估性能或我们的方法和其他可用的已公布方法,使不同方法的研究人员和最终用户了解它们的特性。最后,我们将修改现有的计算机程序PUMA,以处理处理后的RNA-SEQ数据,以便识别哪些基因在条件之间发生变化,哪些具有相似的表达模式,哪些对数据的差异有最大贡献。
英文摘要
After sequencing of the human genome was completed, Scientists were surprised to discover that there are far fewer protein-coding genes than was previously predicted. One reason that an organism as complex as human can be built from a relatively small number of genes is that each gene encodes more than one protein. An intermediate molecule, messenger RNA (mRNA), carries the information from the genome in the cell nucleus to ribosomes which create proteins. These mRNA molecules are also known as transcripts and their full complement is termed the transcriptome. Before they mature these transcripts are edited to form the template for different proteins. This editing process is called splicing and different transcripts that result are called splice variants or isoforms. An additional complexity in the transcriptome is due to the fact that each gene has multiple copies (for example 2 in human, 6 in wheat) and these different copies, called alleles, can be expressed differently under different conditions or in different tissues. The transcriptome is a collection of transcripts which includes all the allele-specific gene isoforms that are expressed in the cell along with other non-coding RNA molecules. Splicing and allele usage are fundamental ways that the function of genes can be modulated in a tissue-specific manner. Therefore developing technologies to accurately measure transcript expression is a necessary step towards understanding and modelling cells and tissues. A recently developed experimental technology called RNA-seq gives unprecedented access to data about the transcriptome. Computational methods are required to interpret these data which are in the form of a list containing millions of short RNA sequence fragments. These fragments are difficult to interpret because, for example, the same fragment could have come from a large number of different gene isoforms. The question is, which one? Computational methods can be used to answer this question and infer the concentration of different gene isoforms in the sample given these data. In this project we will develop a new computational method, implemented in publically available free software, which uses advanced statistical procedures to solve this problem. An important distinguishing feature of the method is the ability to associate inferred concentrations with a degree of uncertainty which captures technical and biological sources of error as well as the inherent difficulty of the problem due to the difficulty of assigning fragments to gene isoforms. We will create benchmark data that allows us to assess the performance or our method and other available published methods, allowing researchers and end-users of different methods to understand their properties. Finally, we will adapt an existing computer program, puma, to work with the processed RNA-seq data in order to identify genes which change between conditions, which have similar expression patterns or which contribute most to the variance in the data.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
Improved variational Bayes inference for transcript expression estimation.
改进了用于转录表达估计的变分贝叶斯推理。
DOI: 10.1515/sagmb-2013-0054
发表时间: 2014
期刊: Statistical applications in genetics and molecular biology
影响因子: 0.9
作者: [Papastamoulis P]
通讯作者: Papastamoulis P
Fast and accurate approximate inference of transcript expression from RNA-seq data
从 RNA-seq 数据快速准确地近似推断转录本表达
DOI: 10.48550/arxiv.1412.5995
发表时间: 2014
期刊:
影响因子: --
作者: [Hensman J]
通讯作者: Hensman J
DOI: 10.1093/bioinformatics/btv483
发表时间: 2015-12-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者: [Hensman J, Papastamoulis P, Glaus P, Honkela A, Rattray M]
通讯作者: Rattray M
DOI: 10.1111/rssc.12213
发表时间: 2018-01
期刊: Journal of the Royal Statistical Society. Series C, Applied statistics
影响因子: --
作者: [Papastamoulis P, Rattray M]
通讯作者: Rattray M
Integrating Capture-HiC with omic time course data to uncover the regulatory interactions modulated by genetic variation in disease
  • 批准号:
    MR/N00017X/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $77.49万
  • 财政年份:
    2015
  • 负责人:
    Magnus Rattray
  • 依托单位:
国内基金
海外基金
企业绩效评价的DEA-Benchmarking方法及动态博弈研究
  • 批准号:
    70571028
  • 项目类别:
    面上项目
  • 资助金额:
    16.5万元
  • 批准年份:
    2005
  • 负责人:
    杨印生
  • 依托单位: