High-throughput protein analysis integrating bioinformatics and experimental assays

High-throughput protein analysis integrating bioinformatics and experimental assays
复制标题

DOI:
10.1093/nar/gkh257
复制
发表时间:
2004-01-01
影响因子:
14.9
通讯作者:
Wiemann, S
Wiemann, S
中科院分区:
生物学2区
文献类型:
--
作者:
del Val, C;Mehrle, A;Wiemann, S

文献摘要

被引文献

相似文献

近年来,大量的转录本信息已被公开,这就需要发展高通量的功能基因组学和蛋白质组学方法来进行分析。这种方法需要适当的数据整合程序和高度自动化,以便从产生的结果中获得最大利益。我们设计了一个自动流水线来分析注释的开放阅读框架(ORF),主要由德国cDNA联盟生产的全长cDNA。将ORF克隆到表达载体中,用于大规模测定,例如确定亚细胞蛋白定位或激酶反应特异性。此外,所有识别的ORF都经过详尽的生物信息学分析,如相似性搜索,蛋白质结构域结构确定和物理化学特征和二级结构的预测,使用各种各样的生物信息学方法与最新的公共数据库(例如PRINTS,BLOCKS,INTERPRO,PROSITE SWISSPROT)相结合。来自实验结果和生物信息学分析的数据被整合并存储在关系数据库(MS SQL Server)中,这使得研究人员可以轻松找到生物学问题的答案,从而加快进一步分析的目标选择。设计的流水线构成了一种新的自动化方法,用于从cDNA的高通量研究中获得和管理相关生物数据,以便系统地识别和表征新基因,以及全面描述编码蛋白质的功能。
The wealth of transcript information that has been made publicly available in recent years requires the development of high-throughput functional genomics and proteomics approaches for its analysis. Such approaches need suitable data integration procedures and a high level of automation in order to gain maximum benefit from the results generated. We have designed an automatic pipeline to analyse annotated open reading frames (ORFs) stemming from full-length cDNAs produced mainly by the German cDNA Consortium. The ORFs are cloned into expression vectors for use in large-scale assays such as the determination of subcellular protein localization or kinase reaction specificity. Additionally, all identified ORFs undergo exhaustive bioinformatic analysis such as similarity searches, protein domain architecture determination and prediction of physicochemical characteristics and secondary structure, using a wide variety of bioinformatic methods in combination with the most up-to-date public databases (e.g. PRINTS, BLOCKS, INTERPRO, PROSITE SWISSPROT). Data from experimental results and from the bioinformatic analysis are integrated and stored in a relational database (MS SQL-Server), which makes it possible for researchers to find answers to biological questions easily, thereby speeding up the selection of targets for further analysis. The designed pipeline constitutes a new automatic approach to obtaining and administrating relevant biological data from high-throughput investigations of cDNAs in order to systematically identify and characterize novel genes, as well as to comprehensively describe the function of the encoded proteins.