Classifying disease outbreak reports using n-grams and semantic features

Classifying disease outbreak reports using n-grams and semantic features
复制标题

DOI:
10.1016/j.ijmedinf.2009.03.010
复制
发表时间:
2009-12-01
影响因子:
4.9
通讯作者:
Collier, Nigel
Collier, Nigel
中科院分区:
医学2区
文献类型:
--
作者:
Conway, Mike;Doan, Son;Collier, Nigel

文献摘要

被引文献

相似文献

引言:本文以BioCaster疾病暴发报告文本挖掘系统为背景,探讨了使用n-gram和语义特征对疾病暴发报告进行分类的好处。这项工作的一个新特征是使用通用语义标记器-USAS标记器-来生成特征。背景:我们概述了这项工作(BioCaster流行病学文本挖掘系统)的应用背景,然后描述了我们的分类实验中使用的实验数据(1000个文档BioCaster语料库)。特征集:在这项工作中使用了三组广泛的特征:基于命名实体的特征,n-gram特征,以及从USAS语义标记器派生的特征。方法:三种标准的机器学习算法-朴素贝叶斯,支持向量机算法,和C4.5决策树算法-用于对实验数据(即BioCaster语料库)进行分类。使用CHI(2)特征选择算法进行特征选择。结果:特征表示法结合朴素贝叶斯算法和特征选择,得到了最高的分类准确率(和F分)。这一结果与基线单字表示法和同一任务之前的工作相比在统计学上具有重要意义。结论:本研究表明,对于疾病暴发报告的分类,词袋、n元语法和语义特征的组合与特征选择相结合,在统计显著水平上提高了分类准确率。(C)2009爱思唯尔爱尔兰有限公司。保留所有权利。
Introduction: This paper explores the benefits of using n-grams and semantic features for the classification of disease outbreak reports, in the context of the BioCaster disease outbreak report text mining system. A novel feature of this work is the use of a general purpose semantic tagger - the USAS tagger - to generate features.Background: We outline the application context for this work (the BioCaster epidemiological text mining system), before going on to describe the experimental data used in our classification experiments (the 1000 document BioCaster corpus).Feature sets: Three broad groups of features are used in this work: Named Entity based features, n-gram features, and features derived from the USAS semantic tagger.Methodology: Three standard machine learning algorithms - Naive Bayes, the Support Vector Machine algorithm, and the C4.5 decision tree algorithm - were used for classifying experimental data (that is, the BioCaster corpus). Feature selection was performed using the chi(2) feature selection algorithm. Standard text classification performance metrics - Accuracy, Precision, Recall, Specificity and F-score - are reported.Results: A feature representation composed of unigrams, bigrams, trigrams and features derived from a semantic tagger, in conjunction with the Naive Bayes algorithm and feature selection yielded the highest classification accuracy (and F-score). This result was statistically significant compared to a baseline unigram representation and to previous work on the same task. However, it was feature selection rather than semantic tagging that contributed most to the improved performance.Conclusion: This study has shown that for the classification of disease outbreak reports, a combination of bag-of-words, n-grams and semantic features, in conjunction with feature selection, increases classification accuracy at a statistically significant level compared to previous work in this domain. (C) 2009 Elsevier Ireland Ltd. All rights reserved.