Heart transcriptome of the bank vole (Myodes glareolus): towards understanding the evolutionary variation in metabolic rate

Heart transcriptome of the bank vole (Myodes glareolus): towards understanding the evolutionary variation in metabolic rate
复制标题

DOI:
10.1186/1471-2164-11-390
复制
发表时间:
2010-06-21
期刊:
影响因子:
4.4
通讯作者:
Radwan, Jacek
Radwan, Jacek
中科院分区:
生物学2区
文献类型:
--
作者:
Babik, Wieslaw;Stuglik, Michal;Radwan, Jacek

文献摘要

被引文献

相似文献

背景:了解适应性变化的遗传基础一直是进化生物学的主要目标。在没有测序基因组的复杂生物体中,使用较长读段测序技术进行从头转录组组装,然后使用短读段进行表达谱分析,可能会在表达水平和编码区序列多态性方面提供适应性变异的全面鉴定。我们在选择用于高代谢和高代谢对照的品系中对库田鼠心脏转录组进行测序和从头组装。结果:一次454钛运行产生了超过百万个读段,这些读段组装成63,581个重叠群。对SwissProt蛋白质数据库和ENSEMBL收集的小鼠转录本的比对分别检测到与11,181和14,051个基因的相似性。根据来自心脏相关基因本体分类的基因和在小鼠心脏中检测到的UniGenes的代表性判断,我们对心脏中表达的基因的检测几乎完成(分别> 95%和几乎90%)。平均而言,我们的序列覆盖了38.7%的转录本长度,编码区的覆盖率(45.0%)明显高于非翻译区(5 'UTR的24.5%和32.7%)。小鼠和银行田鼠之间的非翻译区的序列保守性较低,被认为是部分负责较差的UTR代表。我们的数据可能表明一个广泛的转录从非编码基因组区域,在以前的研究中没有报道的发现在非模式生物的转录组。我们还确定了超过19000个推定的单核苷酸多态性(SNP)。一个更高的分数的SNPs比预期的机会表现出变异频率之间的差异选择regimes.Conclusion:较长的读取和较高的序列产量,每次运行提供的454钛技术相比,前几代焦磷酸测序证明有利于装配的质量。一个几乎完整的代表性基因已知在小鼠心脏中表达的被确定。广泛的基因组资源可用于家鼠,一个适度(20-40百万年)不同的相对的田鼠,使转录完整性的全面评估。本研究中产生的转录序列允许鉴定与选择系分歧相关的候选SNP,并构成有价值的永久资源,形成RNAseq实验的基础,该实验旨在检测基因表达和序列变体水平的适应性变化,这将有助于研究进化分歧的遗传基础。
Background: Understanding the genetic basis of adaptive changes has been a major goal of evolutionary biology. In complex organisms without sequenced genomes, de novo transcriptome assembly using a longer read sequencing technology followed by expression profiling using short reads is likely to provide comprehensive identification of adaptive variation at the expression level and sequence polymorphisms in coding regions. We performed sequencing and de novo assembly of the bank vole heart transcriptome in lines selected for high metabolism and unselected controls.Results: A single 454 Titanium run produced over million reads, which were assembled into 63,581 contigs. Searches against the SwissProt protein database and the ENSEMBL collection of mouse transcripts detected similarity to 11,181 and 14,051 genes, respectively. As judged by the representation of genes from the heart-related Gene Ontology categories and UniGenes detected in the mouse heart, our detection of the genes expressed in the heart was nearly complete (> 95% and almost 90% respectively). On average, 38.7% of the transcript length was covered by our sequences, with notably higher (45.0%) coverage of coding regions than of untranslated regions (24.5% of 5' and 32.7% of 3'UTRs). Lower sequence conservation between mouse and bank vole in untranslated regions was found to be partially responsible for poorer UTR representation. Our data might suggest a widespread transcription from noncoding genomic regions, a finding not reported in previous studies regarding transcriptomes in non-model organisms. We also identified over 19 thousand putative single nucleotide polymorphisms (SNPs). A much higher fraction of the SNPs than expected by chance exhibited variant frequency differences between selection regimes.Conclusion: Longer reads and higher sequence yield per run provided by the 454 Titanium technology in comparison to earlier generations of pyrosequencing proved beneficial for the quality of assembly. An almost full representation of genes known to be expressed in the mouse heart was identified. Usage of the extensive genomic resources available for the house mouse, a moderately (20-40 mln years) divergent relative of the voles, enabled a comprehensive assessment of the transcript completeness. Transcript sequences generated in the present study allowed the identification of candidate SNPs associated with divergence of selection lines and constitute a valuable permanent resource forming a foundation for RNAseq experiments aiming at detection of adaptive changes both at the level of gene expression and sequence variants, that would facilitate studies of the genetic basis of evolutionary divergence.