Mining the Arabidopsis and rice genomes for cyclophilin protein families.
Mining the Arabidopsis and rice genomes for cyclophilin protein families.
复制标题
DOI:
10.1504/ijbra.2009.026421
复制
发表时间:
2009
影响因子:
--
通讯作者:
Moriyama EN
中科院分区:
文献类型:
--
作者:
Opiyo SO;Moriyama EN
Cyclophilins are a family of proteins that possess peptidyl-prolyl isomerase activity. They are present in both eukaryotes and prokaryotes. They are cellular targets of immunosuppressant drugs and involved in a wide variety of functions. The Arabidopsis thaliana genome contains the largest number of cyclophilins. However, the total number of plant cyclophilins available in sequence databases is small compared to that of other organisms. This implies that many cyclophilins are not yet identified in plants. In order to identify cyclophilin candidates from available plant sequence data, we examined alignment-free methods based on partial least squares (PLS) using physico-chemical properties for the mining of single and multiple-domain cyclophilins. PLS with selected descriptors after auto and cross-covariance (ACC) transformation had low false positives compared to PLS with all ACC descriptors. The former PLS classifier also performed better than profile hidden Markov models and PSI-BLAST in identifying cyclophilins from the Arabidopsis and rice genomes.