Parametric Treatment of cDNA Microarray Data
Parametric Treatment of cDNA Microarray Data
复制标题
cDNA 微阵列数据的参数化处理
DOI:
--
复制
发表时间:
2002
期刊:
影响因子:
--
通讯作者:
T. Konishi
中科院分区:
文献类型:
--
作者:
T. Konishi
As each set of microarray data is a ected by variations in experimental conditions, appropriate nor-malizationprocesses are required. Various approaches towards such normalization have been proposedand generally involve adjustments to every pair of the data sets, often between the simultaneouslyhybridizing R and G probes. Chen et al. [1] introduced a model in which each the signal ratio betweenthe probes is normally distributed; the same concept is basically used in the process of the StanfordMicroarray Database [6]. Recently, Yang et al. [4] extended this method by stabilizing the average ofthe signal ratio, which could be biased based on signal intensity. Such non-parametric methods havebeen further improved by introduction of a parameter that also maintains the variance in the signalratios [2, 3]. An alternative and simple parametric method that assumes lognormal distribution ofdata has also been widely employed. Besides its simple ease in performing the calculations, it is alsocapable of determining the data z-scores, a possible common unit for data comparisons. However, mi-croarray data often have a skewed distribution, and this low delity to the distribution model severelylimits the accuracy of the data.In the parametric process, the estimation of background, which is de ned as the constant part ofadditive noise [1], can be a major source of normalization inaccuracies. In most cases, the backgroundis estimated from the area outside the DNA spot of the image data; based on the assumption thatthe background on a tip is uniform. However, as DNA spots and also the intact surface of the tipcan bind free dyes at di erent densities, such estimations are clearly prone to errors. To check thispossibility, an alternative estimation method that is based on signal intensity of control DNA is testedagainst the conventional image-based method. The distributions of both sets of processed data areinvestigated using probability plots.