Informatics strategies for large-scale novel cross-linking analysis

Informatics strategies for large-scale novel cross-linking analysis
复制标题

DOI:
10.1021/pr070035z
复制
发表时间:
2007-09-01
影响因子:
4.4
通讯作者:
Bruce, James E.
Bruce, James E.
中科院分区:
生物学2区
文献类型:
--
作者:
Anderson, Gordon A.;Tolic, Nikola;Bruce, James E.

文献摘要

被引文献

相似文献

生物系统中蛋白质相互作用的检测对当今的技术来说是一个重大挑战。化学交联提供了在复杂系统中赋予新化学键的潜力,从而导致质谱检测到的一组胰蛋白酶肽的质量变化。然而,系统的复杂性和交联产物的异质性阻碍了化学交联广泛用于蛋白质-蛋白质相互作用的大规模鉴定。称为蛋白质相互作用报告基因 (PIR) 的质谱可识别交联剂的开发使得细胞上化学交联实验和产品类型区分成为可能。然而,PIR 实验产生的复杂数据集需要新的信息学能力来进行解释。本手稿详细介绍了我们为开发此类功能所做的努力,并描述了 X-links 程序,该程序允许 PIR 产品类型区分。此外,我们还介绍了 PIR 型实验的蒙特卡罗模拟结果,通过观察到的前体和释放的肽质量为 PIR 产品类型识别提供错误发现率估计。我们的模拟还提供基于精确质量和数据库复杂性的肽识别计算,可以提供肽识别错误发现率的估计。总体而言,计算显示 PIR 产品类型的错误发现率较低,因为随机质量匹配约为 12%,质量测量精度为 10 ppm,并且 100 个肽产生的光谱复杂性。此外,考虑到包含 367 个蛋白质的 Shewanella oneidensis MR-1 的第 1 阶段分析所产生的精简数据库,与整个 Shewanella oneidensis MR-1 蛋白质组相比,预期识别错误发现率估计显着降低。
The detection of protein interactions in biological systems represents a significant challenge for today's technology. Chemical cross-linking provides the potential to impart new chemical bonds in a complex system that result in mass changes in a set of tryptic peptides detected by mass spectrometry. However, system complexity and cross-linking product heterogeneity have precluded widespread chemical cross-linking use for large-scale identification of protein-protein interactions. The development of mass spectrometry identifiable cross-linkers called protein interaction reporters (PIRs) has enabled on-cell chemical cross-linking experiments with product type differentiation. However, the complex datasets resultant from PIR experiments demand new informatics capabilities to allow interpretation. This manuscript details our efforts to develop such capabilities and describes the program X-links, which allows PIR product type differentiation. Furthermore, we also present the results from Monte Carlo simulation of PIR-type experiments to provide false discovery rate estimates for the PIR product type identification through observed precursor and released peptide masses. Our simulations also provide peptide identification calculations based on accurate masses and database complexity that can provide an estimation of false discovery rates for peptide identification. Overall, the calculations show a low rate of false discovery of PIR product types due to random mass matching of approximately 12% with 10 ppm mass measurement accuracy and spectral complexity resulting from 100 peptides. In addition, consideration of a reduced database resulting from stage 1 analysis of Shewanella oneidensis MR-1 containing 367 proteins resulted in a significant reduction of expected identification false discovery rate estimation compared to that from the entire Shewanella oneidensis MR-1 proteome.