Integrative disease classification based on cross-platform microarray data.

Integrative disease classification based on cross-platform microarray data.
复制标题

DOI:
10.1186/1471-2105-10-s1-s25
复制
发表时间:
2009-01-30
期刊:
影响因子:
3
通讯作者:
Zhou XJ
Zhou XJ
中科院分区:
生物学4区
文献类型:
--
作者:
Liu CC;Hu J;Kalakrishnan M;Huang H;Zhou XJ

文献摘要

被引文献

相似文献

疾病分类是微阵列技术的一个重要应用。然而,大多数基于微阵列的分类器只能处理同一研究中产生的数据,因为不同实验室或不同平台产生的微阵列数据由于系统差异而无法直接进行比较。这个问题严重限制了基于微阵列的疾病分类的实际应用。在这项研究中,我们通过整合来自公共微阵列库的大量异质性微阵列数据集来测试疾病分类的可行性。跨平台数据兼容性是通过在数据集内导出表达对数秩比来创建的。然后可以比较数据集之间的对数秩比向量。此外,我们系统地将数据集的文本注释映射到统一医学语言系统(UMLS)中的概念,从而可以定量分析数据集之间的表型“距离”和疾病类别的自动构建。我们设计了一种新的分类方法ManiSVM,它集成了流形数据转换和SVM学习,以利用数据属性。使用留一数据集交叉验证,ManiSVM实现了70.7%的总体准确率(68.6%的精确度和76.9%的召回率),许多疾病类别的准确率超过80%。我们的研究结果不仅证明了集成疾病分类方法的可行性,而且还表明分类精度随着同质训练数据集的数量而增加。因此,整合方法的力量将随着公共存储库中微阵列数据的不断积累而增加。我们的研究表明,自动化疾病诊断可以是一个重要的和有前途的应用程序的大量昂贵的生成,但免费提供,公共微阵列数据。
Disease classification has been an important application of microarray technology. However, most microarray-based classifiers can only handle data generated within the same study, since microarray data generated by different laboratories or with different platforms can not be compared directly due to systematic variations. This issue has severely limited the practical use of microarray-based disease classification. In this study, we tested the feasibility of disease classification by integrating the large amount of heterogeneous microarray datasets from the public microarray repositories. Cross-platform data compatibility is created by deriving expression log-rank ratios within datasets. One may then compare vectors of log-rank ratios across datasets. In addition, we systematically map textual annotations of datasets to concepts in Unified Medical Language System (UMLS), permitting quantitative analysis of the phenotype "distance" between datasets and automated construction of disease classes. We design a new classification approach named ManiSVM, which integrates Manifold data transformation with SVM learning to exploit the data properties. Using the leave one dataset out cross validation, ManiSVM achieved the overall accuracy of 70.7% (68.6% precision and 76.9% recall) with many disease classes achieving the accuracy higher than 80%. Our results not only demonstrated the feasibility of the integrated disease classification approach, but also showed that the classification accuracy increases with the number of homogenous training datasets. Thus, the power of the integrative approach will increase with the continuous accumulation of microarray data in public repositories. Our study shows that automated disease diagnosis can be an important and promising application of the enormous amount of costly to generate, yet freely available, public microarray data.