Multiplatform single-sample estimates of transcriptional activation

Multiplatform single-sample estimates of transcriptional activation
复制标题

DOI:
10.1073/pnas.1305823110
复制
发表时间:
2013-10-29
影响因子:
11.1
通讯作者:
Johnson, W. Evan
Johnson, W. Evan
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Piccolo, Stephen R.;Withers, Michelle R.;Johnson, W. Evan

文献摘要

被引文献

相似文献

在过去的二十年中,许多生物技术平台已被开发用于高通量基因表达谱分析。然而,由于每个平台都受到特定技术偏见的影响,并产生不同的原始数据分布,研究人员在跨平台整合数据时遇到了困难。数据集成对于数据生成联盟、过渡到更新的分析技术的研究人员以及寻求跨实验聚合数据的个人至关重要。我们通过通用表达代码(UPC)方法解决了这一需求,该方法使用考虑基因组碱基组成和目标区域长度的模型来校正平台特定的背景噪声;该方法还使用混合模型来估计基因在特定的分析样本中是否活跃。后者以0到1的比例产生标准化的UPC值,因此无论分析技术如何,都可以一致地解释它们,从而能够以平台无关的方式开发下游分析管道。UPC方法可应用于单通道和双通道表达微阵列以及下一代测序数据(RNA测序)。此外,UPC仅使用来自给定样品内的信息导出-在处理时不需要辅助样品。因此,UPC适用于个性化医疗工作流程,其中样品必须单独处理而不是批量处理。在各种分析和比较中,UPC在大多数情况下与专为微阵列或RNA测序设计的其他方法相比具有优势。计算UPC的软件可在www.bioconductor.org/packages/release/bioc/html/SCAN.UPC.html上免费获得。
Over the past two decades, many biotechnology platforms have been developed for high-throughput gene expression profiling. However, because each platform is subject to technology-specific biases and produces distinct raw-data distributions, researchers have experienced difficulty in integrating data across platforms. Data integration is crucial to data-generating consortiums, researchers transitioning to newer profiling technologies, and individuals seeking to aggregate data across experiments. We address this need with our Universal exPression Code (UPC) approach, which corrects for platform-specific background noise using models that account for the genomic base composition and length of target regions; this approach also uses a mixture model to estimate whether a gene is active in a particular profiling sample. The latter produces standardized UPC values on a zero-to-one scale, so that they can be interpreted consistently, irrespective of profiling technology, thus enabling downstream analysis pipelines to be developed in a platform-agnostic manner. The UPC method can be applied to one- and two-channel expression microarrays and to next-generation sequencing data (RNA sequencing). Furthermore, UPCs are derived using information from within a given sample only-no ancillary samples are required at processing time. Thus, UPCs are suitable for personalized-medicine workflows where samples must be processed individually rather than in batches. In a variety of analyses and comparisons, UPCs perform comparably to other methods designed specifically for microarrays or RNA sequencing in most settings. Software for calculating UPCs is freely available at www.bioconductor.org/packages/release/bioc/html/SCAN.UPC.html.