Construction of Estimation System of Gene Expression Value by Analyzing Microarray Image
Construction of Estimation System of Gene Expression Value by Analyzing Microarray Image
批准号:
15500202
负责人:
TAKEYA Masaru
金额:
$2.37万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
2003
资助国家:
日本
项目状态:
已结题
起止时间:
2003 至 2004
中文摘要
1.基因表达谱通常根据微阵列中斑点的强度值来计算。这种数值变换导致了微阵列图像上其他潜在信息的丢失。我们开发了一个工具来检测每个点上像素的某种统计特征,并编辑了一个包含基因表达和每个点上的统计特征的数据表。2.基因芯片实验涉及大量容易出错的步骤,导致结果数据中存在高水平的噪声。这种噪声降低了基因表达分析的准确性。从cDNA微阵列获得的双点信号的散点图实际上被表示为二维分布,而不是线性分布。对于多个噪声源,偏离理想线性行为本质上反映了随机波动。在此,我们提出了一种用于估计基因芯片上重复数据的基因表达值的技术。对于该估计,一张幻灯片上所有重复数据的分布被建模为概率分布。在散点图中,分布是由正态二维分布的混合构成的,这些正态二维分布表示由于噪声而引起的基因表达值的波动。采用EM算法对建模参数进行估计。使用贝叶斯估计来计算复制数据被噪声移位的概率。从同一天每隔4h从水稻叶片中提取总RNA,用于6种水稻基因芯片分析。这6个数据集被用来测试所提出的技术。根据真值概率对数据集中的基因进行聚类。聚类法成功地发现了水稻昼夜节律调控的候选基因。
英文摘要
1.The gene expression profile is often calculated from intensity value of spot in microarray. This numerical transformation leads to lose the other potential information on microarray image. We developed a tool to detect some kinds of statistical characteristics of pixel on each spot, and edited a data table with gene expression and statistical characteristics on each spot.2.Microarray experiments involve a large number of error-prone steps that lead to a high level of noise in the resulting data. This noise reduces the accuracy of gene expression analysis. Scatter plots of double-spotted signals obtained from cDNA microarrays are actually represented as two-dimensional, rather than linear, distributions. With multi-noise sources, deviations from ideal linear behavior essentially reflect random fluctuations. We herein proposed a technique for estimating gene expression values for duplicated data on cDNA microarrays. For this estimation, distribution of all duplicated data on one slide is modeled as a probability distribution. In the scatter plots, the distribution is constructed from a mixture of normal two-dimensional distributions, which represent fluctuations in gene expression values due to noise. An EM algorithm is used for estimating the modeling parameters. The probability that duplicated data is shifted by noise is calculated using Bayesian estimation. Total RNAs extracted from rice leaves at 4-hour intervals on the same day were used for 6 rice cDNA microarray assays. These 6 data sets were used to test the proposed technique. Genes in the data sets were subjected to clustering based on probability of true value. Clustering successfully identified candidate genes regulated by circadian rhythms in rice.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金