Large-scale optimization-based classification models in medicine and biology

Large-scale optimization-based classification models in medicine and biology
复制标题

DOI:
10.1007/s10439-007-9317-7
复制
发表时间:
2007-06-01
影响因子:
3.8
通讯作者:
Lee, Eva K.
Lee, Eva K.
中科院分区:
工程技术2区
文献类型:
--
作者:
Lee, Eva K.

文献摘要

被引文献

相似文献

我们提出了新的基于优化的分类模型,是通用的,适合于开发大型异构生物和医学数据集的预测规则。我们的预测模型同时结合了(1)对任何数量的不同组进行分类的能力;(2)将异质类型的属性作为输入的能力;(3)消除生物数据中的噪声和错误的高维数据转换;(4)结合约束以限制误分类率的能力,以及保留判断区域,其提供防止过度训练(过度训练往往导致来自所得预测规则的高误分类率)的保护;以及(5)连续多级分类能力,以处理放置在保留判断区域中的数据点。为了说明分类模型和解决方案引擎的功能和灵活性,以及其多组预测能力,描述了预测模型在广泛的生物和医学问题中的应用。应用包括:肿瘤-鳞状疾病类型的鉴别诊断;预测心脏病的存在/不存在;人类癌症中异常CpG岛甲基化的基因组分析和预测;人类肺癌中运动性和形态学数据的判别分析;用于药物递送的超声细胞破裂的预测;肉瘤治疗中肿瘤形状和体积的鉴定;用于预测早期动脉粥样硬化的生物标志物的判别分析;用于早期诊断糖尿病、衰老、黄斑变性和肿瘤转移的天然和血管生成微血管网络的指纹分析;蛋白质定位位点的预测;以及用于土壤类型分类的卫星图像的模式识别。在所有这些应用中,预测模型产生的正确分类率从80%到100%不等。这为将其用作医疗诊断、监测和决策工具提供了动力。
We present novel optimization-based classification models that are general purpose and suitable for developing predictive rules for large heterogeneous biological and medical data sets. Our predictive model simultaneously incorporates (1) the ability to classify any number of distinct groups; (2) the ability to incorporate heterogeneous types of attributes as input; (3) a high-dimensional data transformation that eliminates noise and errors in biological data; (4) the ability to incorporate constraints to limit the rate of misclassification, and a reserved-judgment region that provides a safeguard against over-training (which tends to lead to high misclassification rates from the resulting predictive rule); and (5) successive multi-stage classification capability to handle data points placed in the reserved-judgment region. To illustrate the power and flexibility of the classification model and solution engine, and its multi-group prediction capability, application of the predictive model to a broad class of biological and medical problems is described. Applications include: the differential diagnosis of the type of erythemato-squamous diseases; predicting presence/absence of heart disease; genomic analysis and prediction of aberrant CpG island meythlation in human cancer; discriminant analysis of motility and morphology data in human lung carcinoma; prediction of ultrasonic cell disruption for drug delivery; identification of tumor shape and volume in treatment of sarcoma; discriminant analysis of biomarkers for prediction of early atherosclerois; fingerprinting of native and angiogenic microvascular networks for early diagnosis of diabetes, aging, macular degeneracy and tumor metastasis; prediction of protein localization sites; and pattern recognition of satellite images in classification of soil types. In all these applications, the predictive model yields correct classification rates ranging from 80 to 100%. This provides motivation for pursuing its use as a medical diagnostic, monitoring and decision-making tool.