Exploration, normalization, and genotype calls of high-density oligonucleotide SNP array data

Exploration, normalization, and genotype calls of high-density oligonucleotide SNP array data
复制标题

DOI:
10.1093/biostatistics/kxl042
复制
发表时间:
2007-04-01
期刊:
影响因子:
2.1
通讯作者:
Irizarry, Rafael A.
Irizarry, Rafael A.
中科院分区:
数学2区
文献类型:
--
作者:
Carvalho, Benilton;Bengtsson, Henrik;Irizarry, Rafael A.

文献摘要

被引文献

相似文献

在大多数微阵列技术中,需要许多关键步骤来将原始强度测量值转换为数据分析师、生物学家和临床医生所依赖的数据。这些数据处理,称为预处理,可以影响最终测量的质量。近年来,高通量基因表达检测是微阵列技术最热门的应用领域。对于该应用,各个小组已经证明,相对于该技术的设计者和制造商引入的特设程序,使用现代统计方法可以显著提高基因表达测量的准确度和精确度。目前,微阵列的其他应用越来越受欢迎。在本文中,我们描述了一种预处理方法,用于识别与感兴趣的表型(如疾病)相关的人类基因组特定基因或区域中的DNA序列变体。特别是,我们描述了一种方法,用于预处理的Affytron单核苷酸多态性芯片和获得基因型调用与预处理的数据。我们演示了我们的程序如何使用3个相对较大的研究,包括其中大量的独立调用的数据,以改善现有的方法。所提出的方法在可从Bioconductor获得的包oligo中实现。
In most microarray technologies, a number of critical steps are required to convert raw intensity measurements into the data relied upon by data analysts, biologists, and clinicians. These data manipulations, referred to as preprocessing, can influence the quality of the ultimate measurements. In the last few years, the high-throughput measurement of gene expression is the most popular application of microarray technology. For this application, various groups have demonstrated that the use of modern statistical methodology can substantially improve accuracy and precision of the gene expression measurements, relative to ad hoc procedures introduced by designers and manufacturers of the technology. Currently, other applications of microarrays are becoming more and more popular. In this paper, we describe a preprocessing methodology for a technology designed for the identification of DNA sequence variants in specific genes or regions of the human genome that are associated with phenotypes of interest such as disease. In particular, we describe a methodology useful for preprocessing Affymetrix single-nucleotide polymorphism chips and obtaining genotype calls with the preprocessed data. We demonstrate how our procedure improves existing approaches using data from 3 relatively large studies including the one in which large numbers of independent calls are available. The proposed methods are implemented in the package oligo available from Bioconductor.