'Unknown' proteins and 'orphan' enzymes: the missing half of the engineering parts list--and how to find it.

'Unknown' proteins and 'orphan' enzymes: the missing half of the engineering parts list--and how to find it.
复制标题

DOI:
10.1042/bj20091328
复制
发表时间:
2009-12-14
期刊:
The Biochemical journal
影响因子:
--
通讯作者:
de Crécy-Lagard V
de Crécy-Lagard V
中科院分区:
其他
文献类型:
--
作者:
Hanson AD;Pribat A;Waller JC;de Crécy-Lagard V

文献摘要

被引文献

相似文献

与其他形式的工程一样,代谢工程需要了解目标系统的组成部分(“部件列表”)。缺乏这些知识会损害合理的工程设计和故障原因的诊断;它也给代谢重建的相关领域带来了问题,代谢重建使用细胞的部件列表在计算机上重建其代谢活动。尽管基因组测序取得了惊人的进展,但由于“未知”蛋白质和“孤儿”酶的双重问题,我们试图操纵的大多数生物体的部分列表仍然非常不完整。前者是从基因组序列推导出的所有蛋白质,它们没有已知的功能,而后者是文献中描述的所有酶(并且通常在EC数据库中编目),它们没有相应的基因被报道。在原核生物的基因组中,未知蛋白质约占蛋白质的一半,而在高等植物和动物中,未知蛋白质的数量远远超过这一比例。孤儿酶占EC数据库的三分之一以上。因此,解决“缺失部分列表”问题是后基因组生物学面临的巨大挑战之一,也是发现生命机制新方面的巨大机会。成功将需要一个协调的社区范围内的攻击,持续多年。在这场攻击中,比较基因组学可能是唯一最有效的策略,因为它可以可靠地预测未知蛋白质和孤儿酶基因的功能。此外,由于数据库和相关工具的激增,该系统具有成本效益,而且部署起来越来越简单。
Like other forms of engineering, metabolic engineering requires knowledge of the components (the ‘parts list’) of the target system. Lack of such knowledge impairs both rational engineering design and diagnosis of the reasons for failures; it also poses problems for the related field of metabolic reconstruction, which uses a cell’s parts list to recreate its metabolic activities in silico. Despite spectacular progress in genome sequencing, the parts lists for most organisms that we seek to manipulate remain highly incomplete, due to the dual problem of ‘unknown’ proteins and ‘orphan’ enzymes. The former are all the proteins deduced from genome sequence that have no known function, and the latter are all the enzymes described in the literature (and often catalogued in the EC database) for which no corresponding gene has been reported. Unknown proteins constitute up to about half of the proteins in prokaryotic genomes, and much more than this in higher plants and animals. Orphan enzymes make up more than a third of the EC database. Attacking the ‘missing parts list’ problem is accordingly one of the great challenges for post-genomic biology, and a tremendous opportunity to discover new facets of life’s machinery. Success will require a co-ordinated community-wide attack, sustained over years. In this attack, comparative genomics is probably the single most effective strategy, for it can reliably predict functions for unknown proteins and genes for orphan enzymes. Furthermore, it is cost-efficient and increasingly straightforward to deploy owing to a proliferation of databases and associated tools.