Identification of reference genes for quantitative expression analysis using large-scale RNA-seq data of Arabidopsis thaliana and model crop plants

Identification of reference genes for quantitative expression analysis using large-scale RNA-seq data of Arabidopsis thaliana and model crop plants
复制标题

DOI:
10.1266/ggs.15-00065
复制
发表时间:
2016-04-01
影响因子:
1.1
通讯作者:
Yano, Kentaro
Yano, Kentaro
中科院分区:
生物学4区
文献类型:
--
作者:
Kudo, Toru;Sasaki, Yohei;Yano, Kentaro

文献摘要

被引文献

相似文献

在定量基因表达分析中,为了对结果进行适当的解释,经常使用内参基因作为内控进行规范化。利用微阵列转录组学数据探索优良的新型内参基因,并通过靶向分析评估常用的内参基因。然而,由于特异性检测基因的数量完全取决于微阵列分析中的探针设计,因此使用微阵列数据进行探索可能会错过一些最佳的内参基因选择。最近出现的RNA测序(RNA-seq)为全面探索内参基因提供了理想的资源,因为这种方法能够检测所有表达的基因,原则上甚至包括未知的基因。我们报告了利用来自拟南芥(Arabidopsis thaliana)、大豆(Glycine max)、番茄(Solanum lycopersicum)和水稻(Oryza sativa)等植物的公开RNA-seq数据对内参基因进行全面探索的结果。为了在尽可能广泛的实验条件下选择适合的内参基因,候选基因通过以下四个步骤进行调查:(1)评估每个实验中每个基因的基础表达水平;(2)评价各实验中各基因的表达稳定性;(3)评价实验中各基因的表达稳定性;(4)根据基因稳定表达的实验次数进行排序,选择排名靠前的基因。利用该方法,拟南芥、大豆、番茄和水稻分别获得了13、10、12和21个候选内参基因。微阵列表达数据证实,所提出的内参基因在广泛实验条件下的表达比常用内参基因更稳定。这些新的内参基因将有助于在各种实验条件下进行的实验中分析基因表达谱。
In quantitative gene expression analysis, normalization using a reference gene as an internal control is frequently performed for appropriate interpretation of the results. Efforts have been devoted to exploring superior novel reference genes using microarray transcriptomic data and to evaluating commonly used reference genes by targeting analysis. However, because the number of specifically detectable genes is totally dependent on probe design in the microarray analysis, exploration using microarray data may miss some of the best choices for the reference genes. Recently emerging RNA sequencing (RNA-seq) provides an ideal resource for comprehensive exploration of reference genes since this method is capable of detecting all expressed genes, in principle including even unknown genes. We report the results of a comprehensive exploration of reference genes using public RNA-seq data from plants such as Arabidopsis thaliana (Arabidopsis), Glycine max (soybean), Solanum lycopersicum (tomato) and Oryza sativa (rice). To select reference genes suitable for the broadest experimental conditions possible, candidates were surveyed by the following four steps: (1) evaluation of the basal expression level of each gene in each experiment; (2) evaluation of the expression stability of each gene in each experiment; (3) evaluation of the expression stability of each gene across the experiments; and (4) selection of top-ranked genes, after ranking according to the number of experiments in which the gene was expressed stably. Employing this procedure, 13, 10, 12 and 21 top candidates for reference genes were proposed in Arabidopsis, soybean, tomato and rice, respectively. Microarray expression data confirmed that the expression of the proposed reference genes under broad experimental conditions was more stable than that of commonly used reference genes. These novel reference genes will be useful for analyzing gene expression profiles across experiments carried out under various experimental conditions.