Functional bioinformatics for Arabidopsis thaliana

Functional bioinformatics for Arabidopsis thaliana
复制标题

DOI:
10.1093/bioinformatics/btl051
复制
发表时间:
2006-05-01
期刊:
影响因子:
5.8
通讯作者:
King, RD
King, RD
中科院分区:
生物学3区
文献类型:
--
作者:
Clare, A;Karwath, A;King, RD

文献摘要

被引文献

相似文献

动机:拟南芥的基因组,它有最好的理解植物基因组,仍然有大约三分之一的基因没有功能注释,无论是从MIPS或TAIR。我们已经将我们的数据挖掘预测(DATA MINING PROCESSING,简称DATA)方法应用于预测这些蛋白质序列的功能类的问题。该方法基于使用混合机器学习/数据挖掘方法来识别关于预测功能的序列的生物信息学数据中的模式。我们使用的数据有关的序列,预测的二级结构,预测的结构域,InterPro模式,序列相似性和expressions data.Results:我们预测的功能类的拟南芥基因的比例很高,目前未知的功能。这些预测是可解释的,并具有良好的测试精度。我们详细描述了七个规则产生。
Motivation: The genome of Arabidopsis thaliana, which has the best understood plant genome, still has approximately one-third of its genes with no functional annotation at all from either MIPS or TAIR. We have applied our Data Mining Prediction (DMP) method to the problem of predicting the functional classes of these protein sequences. This method is based on using a hybrid machine-learning/data-mining method to identify patterns in the bioinformatic data about sequences that are predictive of function. We use data about sequence, predicted secondary structure, predicted structural domain, InterPro patterns, sequence similarity profile and expressions data.Results: We predicted the functional class of a high percentage of the Arabidopsis genes with currently unknown function. These predictions are interpretable and have good test accuracies. We describe in detail seven of the rules produced.