Substring selection for biomedical document classification
Substring selection for biomedical document classification
复制标题
DOI:
10.1093/bioinformatics/btl350
复制
发表时间:
2006-09-01
期刊:
影响因子:
5.8
通讯作者:
Vucetic, Slobodan
中科院分区:
文献类型:
--
作者:
Han, Bo;Obradovic, Zoran;Vucetic, Slobodan
Motivation: Attribute selection is a critical step in development of document classification systems. As a standard practice, words are stemmed and the most informative ones are used as attributes in classification. Owing to high complexity of biomedical terminology, general-purpose stemming algorithms are often conservative and could also remove informative stems. This can lead to accuracy reduction, especially when the number of labeled documents is small. To address this issue, we propose an algorithm that omits stemming and, instead, uses the most discriminative substrings as attributes.Results: The approach was tested on five annotated sets of abstracts from iProLINKthat report on the experimental evidence about five types of protein post-translational modifications. The experiments showed that Naive Bayes and support vector machine classifiers perform consistently better[with area under the ROC curve (AUC) accuracy in range 0.92-0.97] when usingthe proposed attribute selection than when using attributes obtained by the Porter stemmer algorithm (AUC in 0.86-0.93 range). The proposed approach is particularly useful when labeled clatasets are small.Contact: vucetic@ist.temple.eduSupplementary Information: The supplementary data are available from www.ist.tempie.edu/PIRsupplement.