Analyzing large-scale proteomics projects with latent semantic indexing

Analyzing large-scale proteomics projects with latent semantic indexing
复制标题

DOI:
10.1021/pr070461k
复制
发表时间:
2008-01-01
影响因子:
4.4
通讯作者:
Hermjakob, Henning
Hermjakob, Henning
中科院分区:
生物学2区
文献类型:
--
作者:
Klie, Sebastian;Martens, Lennart;Hermjakob, Henning

文献摘要

被引文献

相似文献

自从蛋白质组学数据的公共数据库出现以来,来自高通量实验的易于访问的结果一直在稳步积累。特别是几个大规模项目大大增加了社区可获得的身份识别数量。尽管积累了大量信息,但对这些数据进行的成功分析和发表的很少,使这些项目的最终价值远远低于其潜力。已发表的蛋白质组学数据很少被重新分析的一个突出原因在于原始样本收集和随后的数据记录和处理的异质性。为了说明这种异质性的至少一部分可以补偿,我们在这里应用潜在的语义分析的数据贡献的人类蛋白质组组织的血浆蛋白质组计划(HUPO PPP)。有趣的是,尽管HUPO PPP中应用了广泛的仪器和方法,但我们的分析揭示了几种明显的模式,可用于制定优化蛋白质组学项目规划的具体建议以及未来实验中使用的技术选择。从这些结果中可以清楚地看出,通过潜在语义分析等噪声容忍算法对大量公开可用的蛋白质组学数据进行分析具有很大的前景,目前尚未得到充分利用。
Since the advent of public data repositories for proteomics data, readily accessible results from high-throughput experiments have been accumulating steadily. Several large-scale projects in particular have contributed substantially to the amount of identifications available to the community. Despite the considerable body of information amassed, very few successful analyses have been performed and published on this data, leveling off the ultimate value of these projects far below their potential. A prominent reason published proteomics data is seldom reanalyzed lies in the heterogeneous nature of the original sample collection and the subsequent data recording and processing. To illustrate that at least part of this heterogeneity can be compensated for, we here apply a latent semantic analysis to the data contributed by the Human Proteome Organization's Plasma Proteome Project (HUPO PPP). Interestingly, despite the broad spectrum of instruments and methodologies applied in the HUPO PPP, our analysis reveals several obvious patterns that can be used to formulate concrete recommendations for optimizing proteomics project planning as well as the choice of technologies used in future experiments. It is clear from these results that the analysis of large bodies of publicly available proteomics data by noise-tolerant algorithms such as the latent semantic analysis holds great promise and is currently underexploited.