An individualized predictor of health and disease using paired reference and target samples.

An individualized predictor of health and disease using paired reference and target samples.
复制标题

DOI:
10.1186/s12859-016-0889-9
复制
发表时间:
2016-01-22
期刊:
影响因子:
3
通讯作者:
Hero AO
Hero AO
中科院分区:
生物学4区
文献类型:
--
作者:
Liu TY;Burke T;Park LP;Woods CW;Zaas AK;Ginsburg GS;Hero AO

文献摘要

被引文献

相似文献

考虑一下设计一组复杂的生物标志物来预测患者的健康或疾病状态的问题,当一个人可以将他或她当前的测试样本(称为目标样本)与患者先前获得的健康样本(称为参考样本)配对时。与总体平均参考相比,这个参考样本是个性化的。自动预测算法将配对样本相互比较和对比,可以产生新一代的测试面板,与一个人的健康参考进行比较,以提高预测的准确性。本文开发了这样一个个性化的预测器,并说明了将健康参考纳入预测基因表达板设计的附加价值。目的是预测每个受试者的感染状态,例如,既未暴露也未感染,暴露但未感染,感染的急性前阶段,感染的急性阶段,感染的急性后阶段。利用在大规模连续采样的呼吸道病毒挑战研究中收集的基因微阵列数据,我们量化了将一个人的基线参考与他或她的目标样本配对的诊断优势。整个研究由2886个微阵列芯片组成,分析了151名志愿者在4种不同接种方案(HRV, RSV, H1N1, H3N2)下的12,023个基因。我们在这些数据上训练(交叉验证)参考辅助稀疏多类分类器算法,结果表明,对于H3N2队列,纳入受试者的参考样本可以将预测精度提高14%,对于H1N1队列可以提高至少6%。值得注意的是,这些准确性的提高是通过使用更小的基因组来实现的,例如,H3N2减少39%,H1N1减少31%。预测者选择的生物标志物分为两类:1)在人群中目标样本和参考样本之间倾向于差异表达的对比基因;2)在两个样本中保持不变的强化基因,其功能为内务规范基因。这些基因中的许多对所有4种病毒都是共同的,它们在预测因子中的作用阐明了它们在区分宿主免疫反应的不同状态中所起的作用。如果使用合适的数学预测算法,在生物标志物诊断测试中纳入健康参考可能会在生物标志物较少的情况下提高疾病预测的准确性。本文的在线版本(doi:10.1186/s12859-016-0889-9)包含补充材料,可供授权用户使用。
Consider the problem of designing a panel of complex biomarkers to predict a patient’s health or disease state when one can pair his or her current test sample, called a target sample, with the patient’s previously acquired healthy sample, called a reference sample. As contrasted to a population averaged reference this reference sample is individualized. Automated predictor algorithms that compare and contrast the paired samples to each other could result in a new generation of test panels that compare to a person’s healthy reference to enhance predictive accuracy. This paper develops such an individualized predictor and illustrates the added value of including the healthy reference for design of predictive gene expression panels. The objective is to predict each subject’s state of infection, e.g., neither exposed nor infected, exposed but not infected, pre-acute phase of infection, acute phase of infection, post-acute phase of infection. Using gene microarray data collected in a large scale serially sampled respiratory virus challenge study we quantify the diagnostic advantage of pairing a person’s baseline reference with his or her target sample. The full study consists of 2886 microarray chips assaying 12,023 genes of 151 human volunteer subjects under 4 different inoculation regimes (HRV, RSV, H1N1, H3N2). We train (with cross-validation) reference-aided sparse multi-class classifier algorithms on this data to show that inclusion of a subject’s reference sample can improve prediction accuracy by as much as 14 %, for the H3N2 cohort, and by at least 6 %, for the H1N1 cohort. Remarkably, these gains in accuracy are achieved by using smaller panels of genes, e.g., 39 % fewer for H3N2 and 31 % fewer for H1N1. The biomarkers selected by the predictors fall into two categories: 1) contrasting genes that tend to differentially express between target and reference samples over the population; 2) reinforcement genes that remain constant over the two samples, which function as housekeeping normalization genes. Many of these genes are common to all 4 viruses and their roles in the predictor elucidate the function that they play in differentiating the different states of host immune response. If one uses a suitable mathematical prediction algorithm, inclusion of a healthy reference in biomarker diagnostic testing can potentially improve accuracy of disease prediction with fewer biomarkers. The online version of this article (doi:10.1186/s12859-016-0889-9) contains supplementary material, which is available to authorized users.