Automatic document classification of biological literature

Automatic document classification of biological literature
复制标题

DOI:
10.1186/1471-2105-7-370
复制
发表时间:
2006-08-07
期刊:
影响因子:
3
通讯作者:
Sternberg, Paul W.
Sternberg, Paul W.
中科院分区:
生物学4区
文献类型:
--
作者:
Chen, David;Muller, Hans-Michael;Sternberg, Paul W.

文献摘要

被引文献

相似文献

背景:文档分类是许多应用程序中广泛存在的问题,从组织搜索引擎片段到垃圾邮件过滤。我们之前描述过 Textpresso,这是一种生物文献文本挖掘系统,它根据包含生物学兴趣术语的浅层本体来标记全文。该项目利用秀丽隐杆线虫文献语料库的 Textpresso 标记,研究生物文献背景下的文档分类。结果:我们提出了一种两步文本分类算法,对秀丽隐杆线虫论文语料库进行分类。我们的分类方法首先使用支持向量机训练的分类器,然后使用新颖的基于短语的聚类算法。此聚类步骤自动创建人类具有描述性且易于理解的聚类标签。与之前发布的结果(F 值为 0.55 与 0.49)相比,该聚类引擎在标准测试集 (Reuters 21578) 上表现更好,同时生成的聚类描述似乎更有用。 Web 界面允许研究人员快速浏览层次结构并查找属于特定概念的文档。结论:我们已经演示了一种对生物文档进行分类的简单方法,该方法体现了对当前方法的改进。虽然分类结果目前通过人为创建的规则针对秀丽隐杆线虫论文进行了优化,但分类引擎可以适应不同类型的文档。我们通过提供一个网络界面来证明这一点,该界面允许研究人员快速浏览层次结构并查找属于特定概念的文档。
Background: Document classification is a wide-spread problem with many applications, from organizing search engine snippets to spam filtering. We previously described Textpresso, a text-mining system for biological literature, which marks up full text according to a shallow ontology that includes terms of biological interest. This project investigates document classification in the context of biological literature, making use of the Textpresso markup of a corpus of Caenorhabditis elegans literature.Results: We present a two-step text categorization algorithm to classify a corpus of C. elegans papers. Our classification method first uses a support vector machine-trained classifier, followed by a novel, phrase-based clustering algorithm. This clustering step autonomously creates cluster labels that are descriptive and understandable by humans. This clustering engine performed better on a standard test-set (Reuters 21578) compared to previously published results (F-value of 0.55 vs. 0.49), while producing cluster descriptions that appear more useful. A web interface allows researchers to quickly navigate through the hierarchy and look for documents that belong to a specific concept.Conclusion: We have demonstrated a simple method to classify biological documents that embodies an improvement over current methods. While the classification results are currently optimized for Caenorhabditis elegans papers by human-created rules, the classification engine can be adapted to different types of documents. We have demonstrated this by presenting a web interface that allows researchers to quickly navigate through the hierarchy and look for documents that belong to a specific concept.