Automatic assignment of biomedical categories: toward a generic approach

Automatic assignment of biomedical categories: toward a generic approach
复制标题

DOI:
10.1093/bioinformatics/bti783
复制
发表时间:
2006-03-15
期刊:
影响因子:
5.8
通讯作者:
Ruch, P
Ruch, P
中科院分区:
生物学3区
文献类型:
--
作者:
Ruch, P

文献摘要

被引文献

相似文献

动机:我们报告一个通用的文本分类系统,旨在自动分配生物医学类别的任何输入文本的发展。不像通常的自动文本分类系统,它依赖于从大量的训练数据集提取的数据密集型模型,我们的分类器在很大程度上是独立于data-independent.Methods:为了评估我们的方法的鲁棒性,我们测试系统上两个不同的生物医学术语:医学主题词(MeSH)和基因本体论(GO)。我们的轻量级分类器,基于两个排名模块,结合了模式匹配器和向量空间检索引擎,并使用词干和语言驱动的索引单元。结果和结论:结果显示了短语索引对GO和MeSH分类的有效性,但我们观察到该工具的分类能力取决于受控词汇:高等级的精确度范围从MeSH的90%以上到GO的< 20%,为基于检索方法的分类器建立了新的基线。
Motivation: We report on the development of a generic text categorization system designed to automatically assign biomedical categories to any input text. Unlike usual automatic text categorization systems, which rely on data-intensive models extracted from large sets of training data, our categorizer is largely data-independent.Methods: In order to evaluate the robustness of our approach we test the system on two different biomedical terminologies: the Medical Subject Headings (MeSH) and the Gene Ontology (GO). Our lightweight categorizer, based on two ranking modules, combines a pattern matcher and a vector space retrieval engine, and uses both stems and linguistically-motivated indexing units.Results and Conclusion: Results show the effectiveness of phrase indexing for both GO and MeSH categorization, but we observe the categorization power of the tool depends on the controlled vocabulary: precision at high ranks ranges from above 90% for MeSH to < 20% for GO, establishing a new baseline for categorizers based on retrieval methods.