Parameterization of disorder predictors for large-scale applications requiring high specificity by using an extended benchmark dataset.

Parameterization of disorder predictors for large-scale applications requiring high specificity by using an extended benchmark dataset.
复制标题

DOI:
10.1186/1471-2164-11-s1-s15
复制
发表时间:
2010-02-10
期刊:
影响因子:
4.4
通讯作者:
Maurer-Stroh S
Maurer-Stroh S
中科院分区:
生物学2区
文献类型:
--
作者:
Sirota FL;Ooi HS;Gattermayer T;Schneider G;Eisenhaber F;Maurer-Stroh S

文献摘要

被引文献

相似文献

旨在预测蛋白质紊乱的算法在结构和功能基因组学中发挥着重要作用,因为据报道紊乱区域参与重要的细胞过程。因此,不同的小组独立开发了几种具有不同基本原理的疾病预测方法。为了评估它们在自动化工作流程中的可用性,我们有兴趣确定参数设置和阈值选择,在这些参数设置和阈值选择下,这些预测器的性能变得可以直接比较。首先,我们得出了一个新的基准集,该基准集解释了不同类型的紊乱,并补充了针对同一蛋白质集得出的相似数量的顺序注释。我们表明,使用推荐的默认参数,测试的程序正在产生不同水平的特异性和敏感性的广泛预测。我们确定不同预测变量具有相同误报率的设置。我们评估何时可以将一组预测变量一起运行以得出共识或补充预测。这在需要高特异性的蛋白质组应用框架中非常有用,例如在我们的内部序列分析管道和 ANNIE 网络服务器中。这项工作确定了选择无序预测因子的参数设置和阈值,以便在新导出的基准数据集上以所需的特异性水平产生可比较的结果,该数据集同等地考虑了不同长度的有序和无序区域。
Algorithms designed to predict protein disorder play an important role in structural and functional genomics, as disordered regions have been reported to participate in important cellular processes. Consequently, several methods with different underlying principles for disorder prediction have been independently developed by various groups. For assessing their usability in automated workflows, we are interested in identifying parameter settings and threshold selections, under which the performance of these predictors becomes directly comparable. First, we derived a new benchmark set that accounts for different flavours of disorder complemented with a similar amount of order annotation derived for the same protein set. We show that, using the recommended default parameters, the programs tested are producing a wide range of predictions at different levels of specificity and sensitivity. We identify settings, in which the different predictors have the same false positive rate. We assess conditions when sets of predictors can be run together to derive consensus or complementary predictions. This is useful in the framework of proteome-wide applications where high specificity is required such as in our in-house sequence analysis pipeline and the ANNIE webserver. This work identifies parameter settings and thresholds for a selection of disorder predictors to produce comparable results at a desired level of specificity over a newly derived benchmark dataset that accounts equally for ordered and disordered regions of different lengths.