Prediction of thermostability from amino acid attributes by combination of clustering with attribute weighting: a new vista in engineering enzymes.

Prediction of thermostability from amino acid attributes by combination of clustering with attribute weighting: a new vista in engineering enzymes.
复制标题

DOI:
10.1371/journal.pone.0023146
复制
发表时间:
2011
期刊:
影响因子:
3.7
通讯作者:
Ebrahimi M
Ebrahimi M
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Ebrahimi M;Lakizadeh A;Agha-Golzadeh P;Ebrahimie E;Ebrahimi M

文献摘要

参考文献

被引文献

相似文献

热稳定酶的工程正受到越来越多的关注。尤其是造纸、洗涤剂和生物燃料行业,正在寻求使用环保酶代替有毒的氯化学物质。酶通常在低于60°C的温度下起作用,如果暴露在更高的温度下就会变性。相比之下,由于各种结构适应,一小部分酶可以承受更高的温度。了解参与这种适应的蛋白质属性是设计耐热酶的第一步。我们采用各种监督和无监督机器学习算法以及属性加权方法来寻找有助于酶热稳定性的氨基酸组成属性。具体来说,我们比较了两组酶:中稳态酶和热稳态酶。此外,将属性加权与监督和无监督聚类算法相结合,利用氨基酸组成特性对蛋白质热稳定性进行预测和建模。通过多种机器学习算法挖掘大量蛋白质序列(2090),这些算法基于对800多种氨基酸属性的分析,提高了本研究的准确性。此外,这些模型成功地从蛋白质的一级结构预测了热稳定性。结果表明,结合不确定性和相关属性加权算法的期望最大化聚类可以有效地(100%)对热稳定型和介稳定型蛋白质进行分类。70%的加权方法选择谷氨酰胺含量和亲水残基的频率作为最重要的蛋白质属性。在二肽水平上,Asn-Glu的频率是区分中稳定型和热稳定型酶的关键因素。该研究证明了不考虑序列相似性预测热稳定性的可行性,并将为实验室工程热稳定性酶提供基础。
The engineering of thermostable enzymes is receiving increased attention. The paper, detergent, and biofuel industries, in particular, seek to use environmentally friendly enzymes instead of toxic chlorine chemicals. Enzymes typically function at temperatures below 60°C and denature if exposed to higher temperatures. In contrast, a small portion of enzymes can withstand higher temperatures as a result of various structural adaptations. Understanding the protein attributes that are involved in this adaptation is the first step toward engineering thermostable enzymes. We employed various supervised and unsupervised machine learning algorithms as well as attribute weighting approaches to find amino acid composition attributes that contribute to enzyme thermostability. Specifically, we compared two groups of enzymes: mesostable and thermostable enzymes. Furthermore, a combination of attribute weighting with supervised and unsupervised clustering algorithms was used for prediction and modelling of protein thermostability from amino acid composition properties. Mining a large number of protein sequences (2090) through a variety of machine learning algorithms, which were based on the analysis of more than 800 amino acid attributes, increased the accuracy of this study. Moreover, these models were successful in predicting thermostability from the primary structure of proteins. The results showed that expectation maximization clustering in combination with uncertainly and correlation attribute weighting algorithms can effectively (100%) classify thermostable and mesostable proteins. Seventy per cent of the weighting methods selected Gln content and frequency of hydrophilic residues as the most important protein attributes. On the dipeptide level, the frequency of Asn-Glu was the key factor in distinguishing mesostable from thermostable enzymes. This study demonstrates the feasibility of predicting thermostability irrespective of sequence similarity and will serve as a basis for engineering thermostable enzymes in the laboratory.
DOI: 10.1016/j.pep.2005.09.013
发表时间: 2006-03-01
影响因子: 1.6
作者:
Chantasingh, D;Pootanakit, K;Eurwilaichitr, L
通讯作者: Eurwilaichitr, L
DOI: 10.1016/s0167-7799(98)01193-7
发表时间: 1998-08-01
影响因子: 17.3
作者:
Adams, MWW;Kelly, RM
通讯作者: Kelly, RM
DOI: 10.1186/1746-1448-7-1
发表时间: 2011-05-18
期刊: Saline systems
影响因子: --
作者:
Ebrahimie E;Ebrahimi M;Sarvestani NR;Ebrahimi M
通讯作者: Ebrahimi M
DOI: 10.1093/bioinformatics/btq254
发表时间: 2010-07-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Baumgartner, Christian;Lewis, Gregory D.;Gerszten, Robert E.
通讯作者: Gerszten, Robert E.
DOI: 10.1109/tsmcb.2007.895334
发表时间: 2007-08-01
影响因子: --
作者:
Dancey, Darren;Bandar, Zuhair A.;McLean, David
通讯作者: McLean, David