Adaption of the global test idea to proteomics data with missing values

Adaption of the global test idea to proteomics data with missing values
复制标题

DOI:
10.1093/bioinformatics/btu062
复制
发表时间:
2014-05
期刊:
影响因子:
5.8
通讯作者:
K. Jung;H. Dihazi;Asima Bibi;G. H. Dihazi;T. Beissbarth
K. Jung;H. Dihazi;Asima Bibi;G. H. Dihazi;T. Beissbarth
中科院分区:
生物学3区
文献类型:
--
作者:
K. Jung;H. Dihazi;Asima Bibi;G. H. Dihazi;T. Beissbarth

文献摘要

相似文献

动机全局测试程序经常用于基因表达分析,以研究RNA转录物的功能子集与实验组因子之间的关系。然而,这些程序很少用于分析来自其他来源的高通量数据,如蛋白质组表达数据。将全局测试程序从基因组学转移到蛋白质组学数据的主要困难是获得功能注释的更复杂的方式和某些类型的蛋白质组学数据中缺失值的处理。结果我们提出了一个简单的混合线性模型,结合置换过程和缺失值填补进行整体测试的蛋白质组学实验。这种新方法的动机是通过我们目前研究的小鼠实验中的2-D凝胶电泳获得的蛋白质表达数据。一项模拟研究表明,混合模型的功效和检验水平单独会受到数据集中缺失值的影响。缺失值的插补能够纠正某些模拟设置中的偏倚。我们的新方法提供了对与蛋白质集相关的基因本体(GO)术语进行排名的可能性。在特定蛋白质由2-D凝胶上的多个点表示的情况下,通过将这些点也视为蛋白质组也是有帮助的。我们的数据分析指出,缺乏钙网蛋白和与心肌生物学过程相关的蛋白质组之间存在相关性。可用性和实现我们提出的方法包含在R包“RepeatedHighDim '”中,该包已经包含了基因表达数据的全局测试程序。该软件包可从http://cran.r-project.org/检索。联系klaus.jung@ams.med. uni-goettingen.de。
MOTIVATION Global test procedures are frequently used in gene expression analysis to study the relationship between a functional subset of RNA transcripts and an experimental group factor. However, these procedures have been rarely used for the analysis of high-throughput data from other sources, such as proteome expression data. The main difficulties in transferring global test procedures from genomics to proteomics data are the more complicated way of obtaining functional annotations and the handling of missing values in some types of proteomics data. RESULTS We propose a simple mixed linear model in combination with a permutation procedure and missing values imputation to conduct global tests in proteomics experiments. This new approach is motivated by protein expression data obtained by means of 2-D gel electrophoresis within a mouse experiment of our current research. A simulation study yielded that power and testing level of the mixed model alone can be affected by missing values in the dataset. Imputation of missing values was able to correct for a bias in some simulation settings. Our new approach provides the possibility to rank Gene Ontology (GO) terms associated with protein sets. It is also helpful in the case in which a specific protein is represented by multiple spots on a 2-D gel by considering these spots also as a protein set. Analysis of our data points at correlations between the deficiency of the protein 'calreticulin' and protein sets related to biological processes in the heart muscle. AVAILABILITY AND IMPLEMENTATION Our proposed approach is included in the R-package 'RepeatedHighDim', which already contains a global test procedure for gene expression data. The package can be retrieved from http://cran.r-project.org/. CONTACT klaus.jung@ams.med.uni-goettingen.de.