Identification of sample annotation errors in gene expression datasets

Identification of sample annotation errors in gene expression datasets
复制标题

DOI:
10.1007/s00204-015-1632-4
复制
发表时间:
2015-12-01
影响因子:
6.1
通讯作者:
Rahnenfuehrer, Joerg
Rahnenfuehrer, Joerg
中科院分区:
医学2区
文献类型:
--
作者:
Lohr, Miriam;Hellwig, Birte;Rahnenfuehrer, Joerg

文献摘要

被引文献

相似文献

临床注释的人体组织的综合转录组学分析已被广泛应用于肿瘤学、细胞生物学、免疫学和毒理学。在癌症研究中,基于微阵列的基因表达谱已经成功地应用于疾病实体的亚分类、预测治疗反应和识别细胞机制。原始数据的公共可访问性,以及相应的临床病理参数信息,提供了重用以前分析过的数据的机会,并通过组合多个数据集获得统计能力。然而,结果和结论显然取决于现有信息的可靠性。在这里,我们提出了基于基因表达的方法来识别公共转录组数据集中的样本错误注释。样本混淆可以通过区分男性和女性患者样本的分类器来检测。相关分析确定了同一样品中材料的多个测量值。对45个数据集(包括4913名患者)的分析显示,错误的样本注释影响了40%的分析数据集,可能比以前认为的更为普遍。去除错误标记的样本可能会影响某些数据集的统计评估结果。我们的方法可能有助于识别包含大量差异的单个数据集,并且可以常规地纳入临床基因表达数据的统计分析。
The comprehensive transcriptomic analysis of clinically annotated human tissue has found widespread use in oncology, cell biology, immunology, and toxicology. In cancer research, microarray-based gene expression profiling has successfully been applied to subclassify disease entities, predict therapy response, and identify cellular mechanisms. Public accessibility of raw data, together with corresponding information on clinicopathological parameters, offers the opportunity to reuse previously analyzed data and to gain statistical power by combining multiple datasets. However, results and conclusions obviously depend on the reliability of the available information. Here, we propose gene expression-based methods for identifying sample misannotations in public transcriptomic datasets. Sample mix-up can be detected by a classifier that differentiates between samples from male and female patients. Correlation analysis identifies multiple measurements of material from the same sample. The analysis of 45 datasets (including 4913 patients) revealed that erroneous sample annotation, affecting 40 % of the analyzed datasets, may be a more widespread phenomenon than previously thought. Removal of erroneously labelled samples may influence the results of the statistical evaluation in some datasets. Our methods may help to identify individual datasets that contain numerous discrepancies and could be routinely included into the statistical analysis of clinical gene expression data.