Robust detection and genotyping of single feature polymorphisms from gene expression data.

Robust detection and genotyping of single feature polymorphisms from gene expression data.
复制标题

从基因表达数据中对单特征多态性进行稳健检测和基因分型

DOI:
10.1371/journal.pcbi.1000317
复制
发表时间:
2009-03
影响因子:
4.3
通讯作者:
Luo Z
Luo Z
中科院分区:
生物学2区
文献类型:
--
作者:
Wang M;Hu X;Li G;Leach LJ;Potokina E;Druka A;Waugh R;Kearsey MJ;Luo Z

文献摘要

参考文献

被引文献

相似文献

众所周知,Affymetrix 微阵列广泛用于分别通过 RNA 和基因组 DNA 杂交实验预测全基因组基因表达和全基因组遗传多态性。最近有人提出仅使用 RNA 微阵列数据来整合这两个预测。尽管从RNA微阵列数据中检测单特征多态性(SFP)的能力对于已测序和未测序物种的基因组研究具有许多实际意义,但它对为此目的的微阵列基因表达数据的统计建模和分析提出了巨大的挑战。提出了几种从基因表达谱预测 SFP 的方法。然而,它们的性能极易受到基因差异表达的影响。由此预测的 SFP 最终反映的是差异表达基因,而不是真正的序列多态性。为了解决这个问题,我们开发了一种新的统计方法,将转录本与其靶向探针之间的结合亲和力以及测量转录本丰度的参数与 Affymetrix 基因表达数据的完美匹配杂交值分开。我们采用贝叶斯方法来检测 SFP 并在检测到的 SFP 处对分离群体进行基因分型。基于对三个 Affymetrix 微阵列数据集的分析,我们证明,与文献中的竞争对手相比,本方法在检测携带真正序列多态性的 SFP 方面具有显着提高的稳健性和准确性。本文开发的方法将为实验基因组学家提供先进的分析工具,以对其微阵列实验进行适当和有效的分析,并为生物统计学家提供对 Affymetrix 微阵列数据的深刻解释。基因组学的最终目标之一是探索基因组中所有基因的结构和功能变异。高密度寡核苷酸微阵列技术能够分别使用 RNA 和基因组 DNA 样本预测全基因组基因表达和全基因组遗传多态性。最近提出的仅使用 RNA 微阵列数据整合这两种预测的提议在基因组学中具有重大的实际意义。然而,开发一种适当的分析方法来从 RNA 表达数据中检测遗传多态性 (SFP) 是必要且非常具有挑战性的,这些数据本质上与各种生物和技术变异来源相关。本文提出了一种从基因表达数据中检测 SFP 的新统计方法。我们证明,与主流文献中从微阵列数据预测 SFP 的方法相比,新方法对基因差异表达引起的变异明显更加稳健,并且提高了调用具有真正序列多态性的 SFP 的可靠性。检测 SFP 的可预测性得到提高,不仅提高了从微阵列信息评估基因表达的准确性,而且还提供了仅使用一组微阵列数据来整合结构和功能分析的机会。
It is well known that Affymetrix microarrays are widely used to predict genome-wide gene expression and genome-wide genetic polymorphisms from RNA and genomic DNA hybridization experiments, respectively. It has recently been proposed to integrate the two predictions by use of RNA microarray data only. Although the ability to detect single feature polymorphisms (SFPs) from RNA microarray data has many practical implications for genome study in both sequenced and unsequenced species, it raises enormous challenges for statistical modelling and analysis of microarray gene expression data for this objective. Several methods are proposed to predict SFPs from the gene expression profile. However, their performance is highly vulnerable to differential expression of genes. The SFPs thus predicted are eventually a reflection of differentially expressed genes rather than genuine sequence polymorphisms. To address the problem, we developed a novel statistical method to separate the binding affinity between a transcript and its targeting probe and the parameter measuring transcript abundance from perfect-match hybridization values of Affymetrix gene expression data. We implemented a Bayesian approach to detect SFPs and to genotype a segregating population at the detected SFPs. Based on analysis of three Affymetrix microarray datasets, we demonstrated that the present method confers a significantly improved robustness and accuracy in detecting the SFPs that carry genuine sequence polymorphisms when compared to its rivals in the literature. The method developed in this paper will provide experimental genomicists with advanced analytical tools for appropriate and efficient analysis of their microarray experiments and biostatisticians with insightful interpretation of Affymetrix microarray data. One of the ultimate goals of genomics is to explore structural and functional variations of all genes in a genome. High-density oligo-microarray techniques enable prediction of genome-wide gene expression and genome-wide genetic polymorphisms from using RNA and genomic DNA samples, respectively. A recent proposal to integrate the two predictions by use of RNA microarray data alone has great practical implications in genomics. However, it is essential but very challenging to develop an appropriate analytical method for detecting genetic polymorphisms (SFPs) from RNA expression data, which are inherently coupled with various sources of biological and technical variations. This paper presents a novel statistical approach to detect SFPs from gene expression data. We demonstrated that the new method is significantly more robust to variation due to differential expression of genes and improves the reliability of calling SFPs that bear genuine sequence polymorphisms than the other five methods in the mainstream literature on SFP prediction from microarray data. The improved predictability of detecting SFPs not only confers accuracy in evaluating gene expression from microarray information, but also opens up an opportunity to integrate structural and functional analyses by using only one set of microarray data.
DOI: 10.1073/pnas.011404098
发表时间: 2001-01-02
影响因子: 11.1
作者:
Li, C;Wong, WH
通讯作者: Wong, WH
DOI: 10.1101/gr.2850605
发表时间: 2005-02-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Ronald, J;Akey, JM;Kruglyak, L
通讯作者: Kruglyak, L
DOI: 10.1073/pnas.091062498
发表时间: 2001-04-24
影响因子: 11.1
作者:
Tusher, VG;Tibshirani, R;Chu, G
通讯作者: Chu, G
DOI: 10.1534/genetics.106.065292
发表时间: 2007-03-01
期刊: GENETICS
影响因子: 3.3
作者:
Hu, X. H.;Wang, M. H.;Luo, Z. W.
通讯作者: Luo, Z. W.
DOI: 10.1101/gr.5011206
发表时间: 2006-06-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
West, Marilyn A. L.;van Leeuwen, Hans;Michelmore, Richard W.
通讯作者: Michelmore, Richard W.