Adaptive Dimensionality Reduction with Semi-Supervision (AdDReSS): Classifying Multi-Attribute Biomedical Data.

Adaptive Dimensionality Reduction with Semi-Supervision (AdDReSS): Classifying Multi-Attribute Biomedical Data.
复制标题

DOI:
10.1371/journal.pone.0159088
复制
发表时间:
2016
期刊:
影响因子:
3.7
通讯作者:
Madabhushi A
Madabhushi A
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Lee G;Romo Bucheli DE;Madabhushi A

文献摘要

被引文献

相似文献

医学诊断通常是一个多属性问题,需要复杂的工具来分析高维生物医学数据。挖掘这些数据通常会导致两个关键的瓶颈:1)用于表示丰富生物数据的特征的高维性,以及2)由于评估每个研究所需的咨询高度特定的医学专业知识的费用而导致的少量标记的训练数据。目前,我们所知道的没有一种方法试图在降维方法的背景下使用主动学习来改善低维表示的构建。我们提出了我们的新方法,AdDReSS(自适应半监督减少),以证明通过AL在嵌入空间中识别的标记实例较少,需要创建一个更具歧视性的嵌入表示相比,随机选择的实例。我们在前列腺基因表达、卵巢蛋白质组学谱、脑磁共振成像和乳腺组织病理学等广泛领域测试了我们的方法。在这些各种高维生物医学数据集中,每个参数都考虑了100多个观察结果,所有实验的中位数分类准确率显示AdDReSS(88.7%)优于SSAGE,使用随机抽样的SSDR方法(85.5%)和图形嵌入(81.5%)。此外,我们发现通过AdDReSS生成的嵌入在Raghavan效率(学习率的衡量标准)方面比SSAGE平均提高了35.95%。我们的研究结果表明,AdDReSS的价值,提供高维生物医学数据的低维表示,同时实现更高的分类率与较少的标记的例子相比,没有主动学习。
Medical diagnostics is often a multi-attribute problem, necessitating sophisticated tools for analyzing high-dimensional biomedical data. Mining this data often results in two crucial bottlenecks: 1) high dimensionality of features used to represent rich biological data and 2) small amounts of labelled training data due to the expense of consulting highly specific medical expertise necessary to assess each study. Currently, no approach that we are aware of has attempted to use active learning in the context of dimensionality reduction approaches for improving the construction of low dimensional representations. We present our novel methodology, AdDReSS (Adaptive Dimensionality Reduction with Semi-Supervision), to demonstrate that fewer labeled instances identified via AL in embedding space are needed for creating a more discriminative embedding representation compared to randomly selected instances. We tested our methodology on a wide variety of domains ranging from prostate gene expression, ovarian proteomic spectra, brain magnetic resonance imaging, and breast histopathology. Across these various high dimensional biomedical datasets with 100+ observations each and all parameters considered, the median classification accuracy across all experiments showed AdDReSS (88.7%) to outperform SSAGE, a SSDR method using random sampling (85.5%), and Graph Embedding (81.5%). Furthermore, we found that embeddings generated via AdDReSS achieved a mean 35.95% improvement in Raghavan efficiency, a measure of learning rate, over SSAGE. Our results demonstrate the value of AdDReSS to provide low dimensional representations of high dimensional biomedical data while achieving higher classification rates with fewer labelled examples as compared to without active learning.