Mining the Arabidopsis and rice genomes for cyclophilin protein families.

Mining the Arabidopsis and rice genomes for cyclophilin protein families.
复制标题

DOI:
10.1504/ijbra.2009.026421
复制
发表时间:
2009
影响因子:
--
通讯作者:
Moriyama EN
Moriyama EN
中科院分区:
其他
文献类型:
--
作者:
Opiyo SO;Moriyama EN

文献摘要

被引文献

相似文献

亲环素是具有肽基脯氨酰异构酶活性的蛋白质家族。它们存在于真核生物和原核生物中。它们是免疫抑制剂药物的细胞靶点,并参与多种功能。拟南芥基因组中含有最多的亲环蛋白。然而,在序列数据库中可获得的植物亲环素的总数与其他生物相比是小的。这意味着许多亲环素尚未在植物中鉴定。为了从现有的植物序列数据中识别亲环素候选者,我们研究了基于偏最小二乘法(PLS)的无干扰方法,使用物理化学性质来挖掘单域和多域亲环素。与所有ACC描述符的PLS相比,自和互协方差(ACC)变换后的选择描述符的PLS具有较低的假阳性。前者的PLS分类器在识别拟南芥和水稻基因组亲环蛋白方面也优于profile hidden Markov模型和PSI-BLAST。
Cyclophilins are a family of proteins that possess peptidyl-prolyl isomerase activity. They are present in both eukaryotes and prokaryotes. They are cellular targets of immunosuppressant drugs and involved in a wide variety of functions. The Arabidopsis thaliana genome contains the largest number of cyclophilins. However, the total number of plant cyclophilins available in sequence databases is small compared to that of other organisms. This implies that many cyclophilins are not yet identified in plants. In order to identify cyclophilin candidates from available plant sequence data, we examined alignment-free methods based on partial least squares (PLS) using physico-chemical properties for the mining of single and multiple-domain cyclophilins. PLS with selected descriptors after auto and cross-covariance (ACC) transformation had low false positives compared to PLS with all ACC descriptors. The former PLS classifier also performed better than profile hidden Markov models and PSI-BLAST in identifying cyclophilins from the Arabidopsis and rice genomes.