Detecting small plant peptides using SPADA (Small Peptide Alignment Discovery Application).

Detecting small plant peptides using SPADA (Small Peptide Alignment Discovery Application).
复制标题

使用 SPADA(小肽比对发现应用程序)检测小植物肽。

DOI:
10.1186/1471-2105-14-335
复制
发表时间:
2013-11-20
期刊:
影响因子:
3
通讯作者:
Young ND
Young ND
中科院分区:
生物学4区
文献类型:
--
作者:
Zhou P;Silverstein KA;Gao L;Walton JD;Nallu S;Guhlin J;Young ND

文献摘要

参考文献

相似文献

近年来研究表明,植物中编码的小肽基因在植物的发育、繁殖和防御反应中发挥着重要作用。然而,流行的相似性搜索工具和基因预测技术通常无法识别属于这类基因的大多数成员。这在很大程度上是由于家族成员之间的高度序列差异以及实验验证的小肽用作同源性搜索和从头预测的训练集的有限可用性。因此,迫切需要实验和计算研究,以进一步推进小肽的准确预测。我们在这里提出了一个同源性为基础的基因预测程序,以准确地预测在基因组水平的小肽。鉴于高质量的配置文件比对,SPADA识别和注释测试基因组中的几乎所有家族成员,其性能优于所有调查的通用基因预测程序。我们发现许多错误的注释,在目前的拟南芥和苜蓿基因组数据库使用SPADA,其中大部分有RNA-Seq表达的支持。我们还表明SPADA对植物中的其他类型的小分泌肽(例如,自交不亲和蛋白同源物)以及植物界以外的非分泌肽(例如,蘑菇,双孢鹅膏菌中的α-鹅膏毒蛋白毒素基因家族)。SPADA是一个免费的软件工具,可以准确地识别和预测具有一个或两个外显子的短肽的基因结构。SPADA能够将来自轮廓对齐的信息并入模型预测过程中,并利用它来对不同的候选模型进行评分。SPADA在预测小植物肽如富含半胱氨酸的肽家族方面实现了高灵敏度和特异性。研究团体将SPADA系统地应用于其他类别的小肽,将大大提高公共基因组数据库中不同蛋白质家族的基因组注释。
Small peptides encoded as one- or two-exon genes in plants have recently been shown to affect multiple aspects of plant development, reproduction and defense responses. However, popular similarity search tools and gene prediction techniques generally fail to identify most members belonging to this class of genes. This is largely due to the high sequence divergence among family members and the limited availability of experimentally verified small peptides to use as training sets for homology search and ab initio prediction. Consequently, there is an urgent need for both experimental and computational studies in order to further advance the accurate prediction of small peptides. We present here a homology-based gene prediction program to accurately predict small peptides at the genome level. Given a high-quality profile alignment, SPADA identifies and annotates nearly all family members in tested genomes with better performance than all general-purpose gene prediction programs surveyed. We find numerous mis-annotations in the current Arabidopsis thaliana and Medicago truncatula genome databases using SPADA, most of which have RNA-Seq expression support. We also show that SPADA works well on other classes of small secreted peptides in plants (e.g., self-incompatibility protein homologues) as well as non-secreted peptides outside the plant kingdom (e.g., the alpha-amanitin toxin gene family in the mushroom, Amanita bisporigera). SPADA is a free software tool that accurately identifies and predicts the gene structure for short peptides with one or two exons. SPADA is able to incorporate information from profile alignments into the model prediction process and makes use of it to score different candidate models. SPADA achieves high sensitivity and specificity in predicting small plant peptides such as the cysteine-rich peptide families. A systematic application of SPADA to other classes of small peptides by research communities will greatly improve the genome annotation of different protein families in public genome databases.
DOI: 10.1073/pnas.89.22.10915
发表时间: 1992-11-15
影响因子: 11.1
作者:
HENIKOFF, S;HENIKOFF, JG
通讯作者: HENIKOFF, JG
DOI: 10.1093/nar/gkr948
发表时间: 2012-01
影响因子: 14.9
作者:
Hunter S;Jones P;Mitchell A;Apweiler R;Attwood TK;Bateman A;Bernard T;Binns D;Bork P;Burge S;de Castro E;Coggill P;Corbett M;Das U;Daugherty L;Duquenne L;Finn RD;Fraser M;Gough J;Haft D;Hulo N;Kahn D;Kelly E;Letunic I;Lonsdale D;Lopez R;Madera M;Maslen J;McAnulla C;McDowall J;McMenamin C;Mi H;Mutowo-Muellenet P;Mulder N;Natale D;Orengo C;Pesseat S;Punta M;Quinn AF;Rivoire C;Sangrador-Vegas A;Selengut JD;Sigrist CJ;Scheremetjew M;Tate J;Thimmajanarthanan M;Thomas PD;Wu CH;Yeats C;Yong SY
通讯作者: Yong SY
DOI: 10.1186/1741-7007-3-7
发表时间: 2005-03-22
期刊: BMC biology
影响因子: 5.4
作者:
Haas BJ;Wortman JR;Ronning CM;Hannick LI;Smith RK Jr;Maiti R;Chan AP;Yu C;Farzad M;Wu D;White O;Town CD
通讯作者: Town CD
DOI: 10.1038/nature10414
发表时间: 2011-08-28
期刊: Nature
影响因子: 64.8
作者:
Gan X;Stegle O;Behr J;Steffen JG;Drewe P;Hildebrand KL;Lyngsoe R;Schultheiss SJ;Osborne EJ;Sreedharan VT;Kahles A;Bohnert R;Jean G;Derwent P;Kersey P;Belfield EJ;Harberd NP;Kemen E;Toomajian C;Kover PX;Clark RM;Rätsch G;Mott R
通讯作者: Mott R
DOI: 10.1101/gr.5836207
发表时间: 2007-05-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Hanada, Kousuke;Zhang, Xu;Shiu, Shin-Han
通讯作者: Shiu, Shin-Han