Analysis techniques for microarray time-series data

Analysis techniques for microarray time-series data
复制标题

DOI:
10.1089/10665270252935485
复制
发表时间:
2002-01-01
影响因子:
1.7
通讯作者:
Zhi, JZ
Zhi, JZ
中科院分区:
生物学4区
文献类型:
--
作者:
Filkov, V;Skiena, S;Zhi, JZ

文献摘要

被引文献

相似文献

我们解决了酵母基因表达的公开数据集可能存在的局限性。我们通过时间序列分析研究已知监管机构的可预测性,结果表明,在 Cho/Spellman 数据集中,不到 20% 的已知监管对表现出很强的相关性。通过分析已知的监管关系,我们设计了一种边缘检测功能,它比标准相关方法更保真地识别候选监管。我们开发了对粗略时间序列数据集进行综合分析的通用方法。其中包括 1) 主要循环数据集中的自动周期检测方法和 2) 相移循环数据集之间的相位检测。我们展示了如何正确纠正比较不同长度和小字母的序列对之间的相关系数的问题。最后,我们注意到,与汉明距离相比,大小为 2 的字母表上的序列的相关系数可能表现出非常违反直觉的行为。
We address possible limitations of publicly available data sets of yeast gene expression. We study the predictability of known regulators via time-series analysis, and show that less than 20% of known regulatory pairs exhibit strong correlations in the Cho/Spellman data sets. By analyzing known regulatory relationships, we designed an edge detection function which identified candidate regulations with greater fidelity than standard correlation methods. We develop general methods for integrated analysis of coarse time-series data sets. These include 1) methods for automated period detection in a predominately cycling data set and 2) phase detection between phase-shifted cyclic data sets. We show how to properly correct for the problem of comparing correlation coefficients between pairs of sequences of different lengths and small alphabets. Finally, we note that the correlation coefficient of sequences over alphabets of size two can exhibit very counterintuitive behavior when compared with the Hamming distance.