Multi-dimensional classification of biomedical text: toward automated, practical provision of high-utility text to diverse users.

Multi-dimensional classification of biomedical text: toward automated, practical provision of high-utility text to diverse users.
复制标题

DOI:
10.1093/bioinformatics/btn381
复制
发表时间:
2008-09-15
期刊:
影响因子:
5.8
通讯作者:
Wilbur, W. John
Wilbur, W. John
中科院分区:
生物学3区
文献类型:
--
作者:
Shatkay, Hagit;Pan, Fengxia;Rzhetsky, Andrey;Wilbur, W. John

文献摘要

参考文献

被引文献

相似文献

动机:目前生物医学文本挖掘的研究主要是从科学文本中提取特定的信息,为生物学家服务。我们注意到,没有“平均生物学家”的客户端;不同的用户有不同的需求。例如,正如在过去的评估工作(BioCreative,TREC,KDD)中所指出的那样,数据库管理员通常对显示实验证据和方法的句子感兴趣。相反,实验室科学家搜索关于蛋白质的已知信息可能会寻找事实,通常以高置信度陈述。文本挖掘系统可以针对特定的最终用户,并变得更有效,如果系统可以首先识别的科学内容,是用户感兴趣的类型丰富的文本区域,检索文档,有许多这样的区域,并专注于从这些区域的事实提取。在这里,我们研究的能力,自动描述和分类这样的文本。我们最近推出了一个多维的分类和注释计划,开发适用于各种各样的生物医学文件和科学声明,同时旨在支持特定的生物医学检索和提取任务。结果如下:该注释方案被应用到一个大型语料库中,由八个独立的注释者进行控制,其中三个独立的注释者独立地标记每个句子。然后,我们训练和测试机器学习分类器,以根据注释自动对句子片段进行分类。我们在这里讨论这个任务中涉及的问题,并提出一个结果的概述。后者强烈表明,自动标注沿沿着大部分维度是高度可行的,这一新的框架,科学的句子分类是适用于实践。联系方式:shatkay@cs.queensu.ca
Motivation: Much current research in biomedical text mining is concerned with serving biologists by extracting certain information from scientific text. We note that there is no ‘average biologist’ client; different users have distinct needs. For instance, as noted in past evaluation efforts (BioCreative, TREC, KDD) database curators are often interested in sentences showing experimental evidence and methods. Conversely, lab scientists searching for known information about a protein may seek facts, typically stated with high confidence. Text-mining systems can target specific end-users and become more effective, if the system can first identify text regions rich in the type of scientific content that is of interest to the user, retrieve documents that have many such regions, and focus on fact extraction from these regions. Here, we study the ability to characterize and classify such text automatically. We have recently introduced a multi-dimensional categorization and annotation scheme, developed to be applicable to a wide variety of biomedical documents and scientific statements, while intended to support specific biomedical retrieval and extraction tasks. Results: The annotation scheme was applied to a large corpus in a controlled effort by eight independent annotators, where three individual annotators independently tagged each sentence. We then trained and tested machine learning classifiers to automatically categorize sentence fragments based on the annotation. We discuss here the issues involved in this task, and present an overview of the results. The latter strongly suggest that automatic annotation along most of the dimensions is highly feasible, and that this new framework for scientific sentence categorization is applicable in practice. Contact: shatkay@cs.queensu.ca
DOI: 10.1108/eb046814
发表时间: 2006-01-01
影响因子: --
作者:
Porter, M. F.
通讯作者: Porter, M. F.
DOI: 10.1016/j.patcog.2004.03.009
发表时间: 2004-09-01
影响因子: 8
作者:
Boutell, MR;Luo, JB;Brown, CM
通讯作者: Brown, CM
DOI: 10.1093/bioinformatics/bth227
发表时间: 2004-09-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Smith, L;Rindflesch, T;Wilbur, WJ
通讯作者: Wilbur, WJ
DOI: 10.1093/bib/6.3.222
发表时间: 2005-09-01
影响因子: 9.5
作者:
Shatkay, H
通讯作者: Shatkay, H
DOI: 10.1186/1471-2105-7-356
发表时间: 2006-07-25
期刊: BMC bioinformatics
影响因子: 3
作者:
通讯作者: --