Automating document classification for the Immune Epitope Database.

Automating document classification for the Immune Epitope Database.
复制标题

自动化免疫表位数据库的文档分类。

DOI:
10.1186/1471-2105-8-269
复制
发表时间:
2007-07-26
期刊:
影响因子:
3
通讯作者:
Peters, Bjoern
Peters, Bjoern
中科院分区:
生物学4区
文献类型:
--
作者:
Wang, Peng;Morgan, Alexander A.;Zhang, Qing;Sette, Alessandro;Peters, Bjoern

文献摘要

参考文献

被引文献

相似文献

免疫表位数据库包含从科学文献中手动整理的免疫表位信息。与其他知识领域的类似项目一样,在确定哪些文章与此相关方面花费了大量精力。我们在这里报告了我们使用朴素贝叶斯分类器自动化这个过程的经验,这些分类器是在领域专家分类的20,910个摘要上训练的。基本分类器性能的改进是通过a)利用PubMed中存储的摘要本身以外的信息B)应用标准特征选择标准和c)提取例如识别肽序列的域特异性特征模式来实现的。我们已经将分类器实现到策展过程中,以确定摘要是否明显相关,是否明显不相关,或者是否无法进行某些分类,在这种情况下,对摘要进行手动分类。在一个独立的数据集上测试这种分类方案,我们在51.1%的自动分类摘要中实现了95%的灵敏度和特异性。通过实现文本分类,我们在不牺牲人类专家分类的敏感性或特异性的情况下加快了参考选择过程。这项研究既为文本分类工具的用户提供了实用的建议,也为工具开发人员提供了一个可以作为基准的大型数据集。
The Immune Epitope Database contains information on immune epitopes curated manually from the scientific literature. Like similar projects in other knowledge domains, significant effort is spent on identifying which articles are relevant for this purpose. We here report our experience in automating this process using Naïve Bayes classifiers trained on 20,910 abstracts classified by domain experts. Improvements on the basic classifier performance were made by a) utilizing information stored in PubMed beyond the abstract itself b) applying standard feature selection criteria and c) extracting domain specific feature patterns that e.g. identify peptides sequences. We have implemented the classifier into the curation process determining if abstracts are clearly relevant, clearly irrelevant, or if no certain classification can be made, in which case the abstracts are manually classified. Testing this classification scheme on an independent dataset, we achieve 95% sensitivity and specificity in the 51.1% of abstracts that were automatically classified. By implementing text classification, we have sped up the reference selection process without sacrificing sensitivity or specificity of the human expert classification. This study provides both practical recommendations for users of text classification tools, as well as a large dataset which can serve as a benchmark for tool developers.
DOI: 10.1145/183422.183423
发表时间: 1994-07-01
影响因子: 5.6
作者:
APTE, C;DAMERAU, F;WEISS, SM
通讯作者: WEISS, SM
DOI: 10.1108/eb046814
发表时间: 2006-01-01
影响因子: --
作者:
Porter, M. F.
通讯作者: Porter, M. F.
DOI: 10.1186/1471-2105-4-11
发表时间: 2003-03-27
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Donaldson, I;Martin, J;de Bruijn, B;Wolting, C;Lay, V;Tuekam, B;Zhang, SD;Baskin, B;Bader, GD;Michalickova, K;Pawson, T;Hogue, CWV
通讯作者: Hogue, CWV
DOI: 10.1186/1471-2105-7-341
发表时间: 2006-07-12
期刊: BMC bioinformatics
影响因子: 3
作者:
Vita R;Vaughan K;Zarebski L;Salimi N;Fleri W;Grey H;Sathiamurthy M;Mokili J;Bui HH;Bourne PE;Ponomarenko J;de Castro R Jr;Chan RK;Sidney J;Wilson SS;Stewart S;Way S;Peters B;Sette A
通讯作者: Sette A
DOI: 10.1093/nar/gkg056
发表时间: 2003-01-01
影响因子: 14.9
作者:
Bader, GD;Betel, D;Hogue, CWV
通讯作者: Hogue, CWV