HIGHLIGHTING RELATIONSHIPS BETWEEN HETEROGENEOUS BIOLOGICAL DATA THROUGH GRAPHICAL DISPLAYS BASED ON REGULARIZED CANONICAL CORRELATION ANALYSIS

HIGHLIGHTING RELATIONSHIPS BETWEEN HETEROGENEOUS BIOLOGICAL DATA THROUGH GRAPHICAL DISPLAYS BASED ON REGULARIZED CANONICAL CORRELATION ANALYSIS
复制标题

DOI:
10.1142/s0218339009002831
复制
发表时间:
2009-06-01
影响因子:
1.6
通讯作者:
Baccini, A.
Baccini, A.
中科院分区:
生物学4区
文献类型:
--
作者:
Gonzalez, I.;Dejean, S.;Baccini, A.

文献摘要

被引文献

相似文献

高通量技术产生的生物学数据越来越丰富,并引起了许多统计学问题。本文讨论了其中之一,当基因表达数据与其他变量共同观察,目的是突出基因表达和这些其他变量之间的显着关系。探索这些关系的一种相关统计方法是典型相关分析(CCA)。不幸的是,在后基因组数据的背景下,变量(基因表达)的数量通常大于单元(样本)的数量,因此无法直接进行CCA:需要正则化版本。我们对来自两项不同研究的数据集应用了正则化CCA,并表明其解释既证明了之前验证的关系,也证明了新的假设。从第一个数据集(营养基因组学研究),我们产生了有趣的假设转录因子途径可能连接肝脂肪酸和基因表达。从第二个数据集(对NCI-60癌细胞系面板的药物基因组学研究),我们确定了新的ABC转运蛋白候选底物,其相关性通过伴随的几个已知substrate.In结论的识别来说明,使用正则化CCA可能与涉及产生高通量数据的一些和各种生物学实验相关。我们在这里展示了它能够增强从这些相对昂贵的实验中得出的相关结论的范围。
Biological data produced by high throughput technologies are becoming more and more abundant and are arousing many statistical questions. This paper addresses one of them; when gene expression data are jointly observed with other variables with the purpose of highlighting significant relationships between gene expression and these other variables. One relevant statistical method to explore these relationships is Canonical Correlation Analysis (CCA). Unfortunately, in the context of postgenomic data, the number of variables (gene expressions) is usually greater than the number of units (samples) and CCA cannot be directly performed: a regularized version is required.We applied regularized CCA on data sets from two different studies and show that its interpretation evidences both previously validated relationships and new hypothesis. From the first data sets (nutrigenomic study), we generated interesting hypothesis on the transcription factor pathways potentially linking hepatic fatty acids and gene expression. From the second data sets (pharmacogenomic study on the NCI-60 cancer cell line panel), we identified new ABC transporter candidate substrates which relevancy is illustrated by the concomitant identification of several known substrates.In conclusion, the use of regularized CCA is likely to be relevant to a number and a variety of biological experiments involving the generation of high throughput data. We demonstrated here its ability to enhance the range of relevant conclusions that can be drawn from these relatively expensive experiments.