Increasing consistency of disease biomarker prediction across datasets.

Increasing consistency of disease biomarker prediction across datasets.
复制标题

DOI:
10.1371/journal.pone.0091272
复制
发表时间:
2014
期刊:
影响因子:
3.7
通讯作者:
Sealfon SC
Sealfon SC
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Chikina MD;Sealfon SC

文献摘要

参考文献

被引文献

相似文献

人类受试者的微阵列研究通常具有有限的样本量,这阻碍了检测与疾病相关的可靠生物标志物的能力,并激发了对跨研究汇总数据的需求。然而,人类基因表达测量可能会受到许多非随机因素的影响,如遗传学,样品制备和组织异质性。这些因素可能导致相关研究之间缺乏一致性,从而限制了其汇总的实用性。我们表明,它是可行的进行自动校正的个人数据集,以减少这种“潜在变量”的影响(没有事先知道的变量),以这样的方式,数据集解决相同的条件下,显示更好的协议,一旦每个被纠正。我们建立我们的替代变量分析方法的方法,但我们证明,原来的算法是不适合分析的人体组织样本是不同类型的细胞的混合物。我们提出了一个修改SVA是至关重要的,以获得改善协议,我们观察到。我们在多发性硬化症数据汇编上开发了我们的方法,并在帕金森病数据集的独立汇编上验证了该方法。在这两种情况下,我们表明我们的方法能够提高不同研究设计、平台和组织的一致性。这种方法有可能广泛适用于任何缺乏研究间协议的领域。
Microarray studies with human subjects often have limited sample sizes which hampers the ability to detect reliable biomarkers associated with disease and motivates the need to aggregate data across studies. However, human gene expression measurements may be influenced by many non-random factors such as genetics, sample preparations, and tissue heterogeneity. These factors can contribute to a lack of agreement among related studies, limiting the utility of their aggregation. We show that it is feasible to carry out an automatic correction of individual datasets to reduce the effect of such ‘latent variables’ (without prior knowledge of the variables) in such a way that datasets addressing the same condition show better agreement once each is corrected. We build our approach on the method of surrogate variable analysis but we demonstrate that the original algorithm is unsuitable for the analysis of human tissue samples that are mixtures of different cell types. We propose a modification to SVA that is crucial to obtaining the improvement in agreement that we observe. We develop our method on a compendium of multiple sclerosis data and verify it on an independent compendium of Parkinson's disease datasets. In both cases, we show that our method is able to improve agreement across varying study designs, platforms, and tissues. This approach has the potential for wide applicability to any field where lack of inter-study agreement has been a concern.
DOI: 10.1093/bioinformatics/bts022
发表时间: 2012-03-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Chakraborty, Sutirtha;Datta, Somnath;Datta, Susmita
通讯作者: Datta, Susmita
DOI: 10.1111/j.1365-2753.1995.tb00005.x
发表时间: 1995-09-01
影响因子: 2.4
作者:
Eysenck, H J
通讯作者: Eysenck, H J
通过替代变量分析捕获基因表达研究中的异质性。
DOI: 10.1371/journal.pgen.0030161
发表时间: 2007-09
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Leek, Jeffrey T.;Storey, John D.
通讯作者: Storey, John D.
使用基于途径的方法对微阵列数据进行荟萃分析确定了人类外周血单核细胞中全身性红斑狼疮的37基因表达签名。
DOI: 10.1186/1741-7015-9-65
发表时间: 2011-05-30
期刊: BMC medicine
影响因子: 9.3
作者:
Arasappan D;Tong W;Mummaneni P;Fang H;Amur S
通讯作者: Amur S
来自多个微阵列实验的基因表达数据的荟萃分析的潜在变量方法。
DOI: 10.1186/1471-2105-8-364
发表时间: 2007-09-27
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Choi, Hyungwon;Shen, Ronglai;Chinnaiyan, Arul M;Ghosh, Debashis
通讯作者: Ghosh, Debashis