Normalization and analysis of residual variation in two-dimensional gel electrophoresis for quantitative differential proteomics

Normalization and analysis of residual variation in two-dimensional gel electrophoresis for quantitative differential proteomics
复制标题

DOI:
10.1002/pmic.200401003
复制
发表时间:
2005-04-01
期刊:
影响因子:
3.4
通讯作者:
Arthur, JM
Arthur, JM
中科院分区:
生物学3区
文献类型:
--
作者:
Almeida, JS;Stanislaus, R;Arthur, JM

文献摘要

被引文献

相似文献

虽然二维凝胶电泳(2-DE)一直是筛选蛋白质组的常用实验方法,但其可重复性很少借助定量误差模型进行分析。残差分布模型的缺乏可用于分配差异表达的可能性,这反映了在处理斑点强度变异性和不同凝胶中对同一斑点的不确定识别的综合影响方面的困难。在本报告中,我们分析了在两个不同发育阶段的鸡胚心脏样品的一系列四个三倍二维凝胶,以产生这样的残差分布模型。为了实现这一参考误差模型,必须建立一致光斑强度归一化的非参数过程,本文也报道了这一过程。除了由于各种来源导致的归一化强度的可变性外,观察到重复之间的残余变异由于无法识别斑点本身(凝胶比对)而变得更加复杂。混合效应反映在残差的双峰密度分布中。适应这种分布的全局误差模型的提取是通过机器学习经验实现的,特别是通过自举人工神经网络。所描述的模型被用来为任意2-DE凝胶中观察到的变化分配置信度值,以量化蛋白质斑点的过表达和低表达程度。
Although two-dimensional gel electrophoresis (2-DE) has long been a favorite experimental method to screen proteomes, its reproducibility is seldom analyzed with the assistance of quantitative error models. The lack of models of residual distributions that can be used to assign likelihood to differential expression reflects the difficulty in tackling the combined effect of variability in spot intensity and uncertain recognition of the same spot in different gels. In this report we have analyzed a series of four triplicate two-dimensional gels of chicken embryo heart samples at two distinct development stages to produce such a model of residual distribution. In order to achieve this reference error model, a nonparametric procedure for consistent spot intensity normalization had to be established, and is also reported here. In addition to variability in normalized intensity due to various sources, the residual variation between replicates was observed to be compounded by failure to identify the spot itself (gel alignment). The mixed effect is reflected by variably skewed bimodal density distributions of residuals. The extraction of a global error model that accommodated such distribution was achieved empirically by machine learning, specifically by bootstrapped artificial neural networks. The model described is being used to assign confidence values to observed variations in arbitrary 2-DE gels in order to quantify the degree of over-expression and under-expression of protein spots.