Multiple feature construction for effective biomarker identification and classification using genetic programming

Multiple feature construction for effective biomarker identification and classification using genetic programming
复制标题

使用遗传编程进行有效生物标志物识别和分类的多特征构建

DOI:
--
复制
发表时间:
2014
期刊:
Annual Conference on Genetic and Evolutionary Computation
影响因子:
--
通讯作者:
Bing Xue
Bing Xue
中科院分区:
--
文献类型:
--
作者:
Soha Ahmed;Mengjie Zhang;Lifeng Peng;Bing Xue

文献摘要

参考文献

被引文献

相似文献

生物标志物鉴定,即检测表明两个或多个类别之间差异的特征,是组学科学的重要任务。质谱(MS)提供蛋白质组学和代谢组学数据的高通量分析。质谱数据集的特征数量远远超过样品的数量,使得生物标志物的鉴定极其困难。特征构造可以通过将原始特征转换为数量较少的高级特征来提供解决这一问题的方法。本文研究了利用遗传规划(GP)构建多特征的方法对生物标志物进行鉴定和质谱数据分类。本文采用嵌入的方法构造了GP的多个特征,其中使用Fisher准则和p值来度量类之间的判别信息。这从二元和多类质谱数据集的低级特征中产生非线性高级特征。同时,使用7种不同的分类器来测试所构建特征的有效性。提出的GP方法在8个不同的质谱数据集上进行了测试。结果表明,在大多数情况下,GP方法构造的高级特征比原始特征集和低级选择的特征更有效地提高了分类性能。此外,新方法在生物标志物的检出率方面表现出优越的性能。
Biomarker identification, i.e., detecting the features that indicate differences between two or more classes, is an important task in omics sciences. Mass spectrometry (MS) provide a high throughput analysis of proteomic and metabolomic data. The number of features of the MS data sets far exceeds the number of samples, making biomarker identification extremely difficult. Feature construction can provide a means for solving this problem by transforming the original features to a smaller number of high-level features. This paper investigates the construction of multiple features using genetic programming (GP) for biomarker identification and classification of mass spectrometry data. In this paper, multiple features are constructed using GP by adopting an embedded approach in which Fisher criterion and p-values are used to measure the discriminating information between the classes. This produces nonlinear high-level features from the low-level features for both binary and multi-class mass spectrometry data sets. Meanwhile, seven different classifiers are used to test the effectiveness of the constructed features. The proposed GP method is tested on eight different mass spectrometry data sets. The results show that the high-level features constructed by the GP method are effective in improving the classification performance in most cases over the original set of features and the low-level selected features. In addition, the new method shows superior performance in terms of biomarker detection rate.
DOI: 10.1016/s1535-6108(03)00309-x
发表时间: 2003-12-01
期刊: CANCER CELL
影响因子: 50.3
作者:
Hingorani, SR;Petricoin, EF;Tuveson, DA
通讯作者: Tuveson, DA
DOI: 10.1007/978-1-60327-194-3_11
发表时间: 2010-01-01
期刊: BIOINFORMATICS METHODS IN CLINICAL RESEARCH
影响因子: --
作者:
Datta, Susmita;Pihur, Vasyl
通讯作者: Pihur, Vasyl
DOI: 10.1021/ac051437y
发表时间: 2006-02-01
影响因子: 7.4
作者:
Smith, CA;Want, EJ;Siuzdak, G
通讯作者: Siuzdak, G
DOI: 10.1021/pr0705237
发表时间: 2008-02-01
影响因子: 4.4
作者:
Ressom, Habtorn W.;Varghese, Rency S.;Goldman, Radoslav
通讯作者: Goldman, Radoslav