PCDq: human protein complex database with quality index which summarizes different levels of evidences of protein complexes predicted from h-invitational protein-protein interactions integrative dataset.

PCDq: human protein complex database with quality index which summarizes different levels of evidences of protein complexes predicted from h-invitational protein-protein interactions integrative dataset.
复制标题

DOI:
10.1186/1752-0509-6-s2-s7
复制
发表时间:
2012
影响因子:
--
通讯作者:
Imanishi T
Imanishi T
中科院分区:
生物2区
文献类型:
--
作者:
Kikugawa S;Nishikata K;Murakami K;Sato Y;Suzuki M;Altaf-Ul-Amin M;Kanaya S;Imanishi T

文献摘要

被引文献

相似文献

蛋白质与复合体中的其他蛋白质或生物分子相互作用,以执行细胞功能。现有的蛋白质-蛋白质相互作用(PPI)数据库和蛋白质复合体数据库没有组织起来提供蛋白质复合体信息或促进新亚基的发现。PPI的数据整合主要集中在蛋白质复合体、亚单位及其功能上。预测的候选复合体或亚单位对实验生物学家也很重要。基于综合的PPI数据和文献,我们开发了一个具有复杂质量指数(PCDq)的人类蛋白质复合体数据库,其中包括已知和预测的复合体及其亚基。我们整合了六个PPI数据(BIND、DIP、MINT、HPRD、INTERNAL和GNP_Y2H),并通过在PPI网络中找到密集连接的区域来预测人类蛋白质复合体。根据文献对它们进行整理,以补充缺失的蛋白质并合并一些复合体,结果得到1,264个复合体,包括32,198个PPI的9,268个蛋白质。每个亚基的证据水平被指定为分类变量。这表明它是否是一个已知的亚基,特定的功能可以从序列或网络分析中推断出来。为了总结复合体中所有亚单位的类别,我们设计了一个复合体质量指数(CQI),并将其分配给每个复合体。我们检查了基因本体论(GO)术语在复合体的蛋白质亚单位之间的一致性比例。接下来,我们比较了相应基因的表达谱,发现在较大的复合体中,许多蛋白质倾向于在转录水平上协同表达。对复合体中重复基因的比例进行了评估。最后,我们确定了78个假设的蛋白质,这些蛋白质被注释为82个复合体的亚基,其中包括已知的复合体。在这些假想的蛋白质中,在我们的预测之后,有四个被报告为指定的蛋白质复合体的实际亚基。我们构建了一个新的蛋白质复合体数据库PCDq,包括预测和精选的人类蛋白质复合体。CQI是实验证实的有关蛋白质复合体和亚单位的有用信息来源。预测的蛋白质复合体可以提供关于假想蛋白质的功能线索。您可以在http://h-invitational.jp/hinv/pcdq/.上免费获得PCDQ
Proteins interact with other proteins or biomolecules in complexes to perform cellular functions. Existing protein-protein interaction (PPI) databases and protein complex databases for human proteins are not organized to provide protein complex information or facilitate the discovery of novel subunits. Data integration of PPIs focused specifically on protein complexes, subunits, and their functions. Predicted candidate complexes or subunits are also important for experimental biologists. Based on integrated PPI data and literature, we have developed a human protein complex database with a complex quality index (PCDq), which includes both known and predicted complexes and subunits. We integrated six PPI data (BIND, DIP, MINT, HPRD, IntAct, and GNP_Y2H), and predicted human protein complexes by finding densely connected regions in the PPI networks. They were curated with the literature so that missing proteins were complemented and some complexes were merged, resulting in 1,264 complexes comprising 9,268 proteins with 32,198 PPIs. The evidence level of each subunit was assigned as a categorical variable. This indicated whether it was a known subunit, and a specific function was inferable from sequence or network analysis. To summarize the categories of all the subunits in a complex, we devised a complex quality index (CQI) and assigned it to each complex. We examined the proportion of consistency of Gene Ontology (GO) terms among protein subunits of a complex. Next, we compared the expression profiles of the corresponding genes and found that many proteins in larger complexes tend to be expressed cooperatively at the transcript level. The proportion of duplicated genes in a complex was evaluated. Finally, we identified 78 hypothetical proteins that were annotated as subunits of 82 complexes, which included known complexes. Of these hypothetical proteins, after our prediction had been made, four were reported to be actual subunits of the assigned protein complexes. We constructed a new protein complex database PCDq including both predicted and curated human protein complexes. CQI is a useful source of experimentally confirmed information about protein complexes and subunits. The predicted protein complexes can provide functional clues about hypothetical proteins. PCDq is freely available at http://h-invitational.jp/hinv/pcdq/.