The use of MPSS for whole-genome transcriptional analysis in Arabidopsis

The use of MPSS for whole-genome transcriptional analysis in Arabidopsis
复制标题

DOI:
10.1101/gr.2275604
复制
发表时间:
2004-08-01
期刊:
影响因子:
7
通讯作者:
Decola, S
Decola, S
中科院分区:
生物学1区
文献类型:
--
作者:
Meyers, BC;Tej, SS;Decola, S

文献摘要

被引文献

相似文献

我们已经产生了36,991,173个代表模式植物拟南芥转录本的17碱基序列“签名”。这些数据通过大规模并行签名测序(MPSS)从14个文库中获得,并且包括268,132个不同的序列。用20个碱基的签名也获得了可比较的数据。我们开发了一种方法来处理这些数据,并比较这些签名注释拟南芥基因组。作为该程序的一部分,从拟南芥基因组中提取858,019个潜在或“基因组”标签,并基于标签相对于注释基因的位置和方向进行分类。基因组和表达的标签的比较匹配67,735个预测来自不同转录物并以显著水平表达的标签。表达的标签来源于29,084个注释基因中至少19,088个的有义链。基因组和表达特征的比较表明,类似于7.7%的基因组特征在表达数据中代表性不足。这些基因组签名包含20个四碱基单词中的一个,这些单词始终与MPSS丰度降低相关。超过89%的表达签名丰度的总和与拟南芥基因组匹配,并且预测在高丰度中发现的许多不匹配的签名与先前未表征的转录本匹配。
We have generated 36,991,173 17-base Sequence "signatures" representing transcripts from the model plant Arabidopsis. These data were derived by massively parallel signature Sequencing (MPSS) from 14 libraries and comprised 268,132 distinct Sequences. Comparable data were also obtained with 20-base signatures. We developed a method for handling these data and for comparing these signatures to the annotated Arabidopsis genome. As part of this procedure, 858,019 potential or "genomic" signatures were extracted from the Arabidopsis genome and classified based on the position and orientation of the signatures relative to annotated genes. A comparison of genomic and expressed signatures matched 67,735 signatures predicted to be derived from distinct transcripts and expressed at significant levels. Expressed signatures were derived from the sense strand of at least 19,088 of 29,084 annotated genes. A comparison of the genomic and expression signatures demonstrated that similar to7.7% of genomic signatures were underrepresented in the expression data. These genomic signatures contained one of 20 four-base words that were consistently associated with reduced MPSS abundances. More than 89% of the sum of the expressed signature abundances matched the Arabidopsis genome, and many of the unmatched signatures found in high abundances were predicted to match to previously uncharacterized transcripts.