Unbiased descriptor and parameter selection confirms the potential of proteochemometric modelling.

Unbiased descriptor and parameter selection confirms the potential of proteochemometric modelling.
复制标题

DOI:
10.1186/1471-2105-6-50
复制
发表时间:
2005-03-10
期刊:
影响因子:
3
通讯作者:
Gustafsson MG
Gustafsson MG
中科院分区:
生物学4区
文献类型:
--
作者:
Freyhult E;Prusis P;Lapinsh M;Wikberg JE;Moulton V;Gustafsson MG

文献摘要

参考文献

被引文献

相似文献

蛋白质化学计量学是一种新的方法,它可以直接从真实的相互作用测量数据预测蛋白质的功能,而不需要三维结构信息。几个报道的配体-受体相互作用的蛋白化学计量学模型已经产生了重要的见解,各种形式的生物分子相互作用。蛋白质化学计量模型是预测配体和蛋白质的特征的特定组合的结合亲和力的多变量回归模型。虽然蛋白化学计量模型已经在各种研究中提供了有趣的结果,但尚未对其平均预测能力进行详细的统计评估。特别是,迄今为止进行的可变子集选择一直依赖于使用所有可用的例子,在微阵列基因表达数据分析中也遇到这种情况。一种无偏的蛋白质化学计量学模型的预测能力的评价方法的实施和应用它的两个最大的蛋白质化学计量学数据集尚未报告的结果。一个双交叉验证循环过程是用来估计一个给定的设计方法的预期性能。无偏的性能估计(P2),我们认为得到的数据集确认,适当设计的单一蛋白化学计量模型具有有用的预测能力,但基于交叉验证的标准设计可能会产生相当有限的性能模型。结果还表明,不同的商业软件包采用的蛋白质化学计量模型的设计可能会产生非常不同的,因此误导的性能估计。此外,在双CV循环中获得的模型中的差异表明,当数据集小时,单个蛋白化学计量模型的详细化学解释是不确定的。所采用的双CV循环提供关于给定蛋白质化学计量建模过程的无偏性能估计,使得可以识别蛋白质化学计量设计不产生有用的预测模型的情况。单蛋白化学计量模型的化学解释是不确定的,而应该是基于所有的模型选择在这里采用的双CV循环。
Proteochemometrics is a new methodology that allows prediction of protein function directly from real interaction measurement data without the need of 3D structure information. Several reported proteochemometric models of ligand-receptor interactions have already yielded significant insights into various forms of bio-molecular interactions. The proteochemometric models are multivariate regression models that predict binding affinity for a particular combination of features of the ligand and protein. Although proteochemometric models have already offered interesting results in various studies, no detailed statistical evaluation of their average predictive power has been performed. In particular, variable subset selection performed to date has always relied on using all available examples, a situation also encountered in microarray gene expression data analysis. A methodology for an unbiased evaluation of the predictive power of proteochemometric models was implemented and results from applying it to two of the largest proteochemometric data sets yet reported are presented. A double cross-validation loop procedure is used to estimate the expected performance of a given design method. The unbiased performance estimates (P2) obtained for the data sets that we consider confirm that properly designed single proteochemometric models have useful predictive power, but that a standard design based on cross validation may yield models with quite limited performance. The results also show that different commercial software packages employed for the design of proteochemometric models may yield very different and therefore misleading performance estimates. In addition, the differences in the models obtained in the double CV loop indicate that detailed chemical interpretation of a single proteochemometric model is uncertain when data sets are small. The double CV loop employed offer unbiased performance estimates about a given proteochemometric modelling procedure, making it possible to identify cases where the proteochemometric design does not result in useful predictive models. Chemical interpretations of single proteochemometric models are uncertain and should instead be based on all the models selected in the double CV loop employed here.
DOI: 10.1021/jm00007a003
发表时间: 1995-03-31
影响因子: 7.3
作者:
CHO, SJ;TROPSHA, A
通讯作者: TROPSHA, A
DOI: 10.1124/mol.61.6.1465
发表时间: 2002-06-01
影响因子: 3.6
作者:
Lapinsh, M;Prusis, P;Wikberg, JES
通讯作者: Wikberg, JES
DOI: 10.1093/protein/15.4.305
发表时间: 2002-04-01
期刊: PROTEIN ENGINEERING
影响因子: --
作者:
Prusis, P;Lundstedt, T;Wikberg, JES
通讯作者: Wikberg, JES
DOI: 10.1021/bi972733a
发表时间: 1998-04-21
期刊: BIOCHEMISTRY
影响因子: 2.9
作者:
Hamaguchi, N;True, TA;Jeffs, PW
通讯作者: Jeffs, PW