On filtering false positive transmembrane protein predictions

On filtering false positive transmembrane protein predictions
复制标题

DOI:
10.1093/protein/15.9.745
复制
发表时间:
2002-09-01
期刊:
PROTEIN ENGINEERING
影响因子:
--
通讯作者:
Simon, I
Simon, I
中科院分区:
其他
文献类型:
--
作者:
Cserzö, M;Eisenhaber, F;Simon, I

文献摘要

被引文献

相似文献

虽然螺旋跨膜(TM)区域预测工具对真实的完整膜蛋白获得了高(90%)的成功率,但它们在已知的非跨膜查询序列中产生了相当数量的假阳性命中。我们提出了一种改进的密集对齐面(DAS)方法,大大降低了误报错误率。本质上,在第二步中,将包括可能的跨膜区域的序列与记录的跨膜蛋白序列文库中的TM片段进行比较。在本试验中,如果查询序列相对于记录的含有TM片段序列的文库的性能低于经验阈值,则将其归类为非跨膜蛋白。可信TM区域命中的误报预测概率用E值表示。改进的DAS方法,即DAS-TM Filter算法,对TM片段具有不变的高灵敏度(类似于在128个已记录的跨膜蛋白学习集中检测到的95%)。同时,对526个已知3D结构的非冗余可溶性蛋白质组的选择性测量接近99%,这主要是因为DAS-TM过滤算法消除了大量错误预测的单膜通过蛋白质。
While helical transmembrane (TM) region prediction tools achieve high (>90%) success rates for real integral membrane proteins, they produce a considerable number of false positive hits in sequences of known nontransmembrane queries. We propose a modification of the dense alignment surface (DAS) method that achieves a substantial decrease in the false positive error rate. Essentially, a sequence that includes possible transmembrane regions is compared in a second step with TM segments in a sequence library of documented transmembrane proteins. If the performance of the query sequence against the library of documented TM segment-containing sequences in this test is lower than an empirical threshold, it is classified as a non-transmembrane protein. The probability of false positive prediction for trusted TM region hits is expressed in terms of E-values. The modified DAS method, the DAS-TMfilter algorithm, has an unchanged high sensitivity for TM segments (similar to95% detected in a learning set of 128 documented transmembrane proteins). At the same time, the selectivity measured over a non-redundant set of 526 soluble proteins with known 3D structure is similar to99%, mainly because a large number of falsely predicted single membrane-pass proteins are eliminated by the DAS-TMfilter algorithm.