A comparison and user-based evaluation of models of textual information structure in the context of cancer risk assessment.

A comparison and user-based evaluation of models of textual information structure in the context of cancer risk assessment.
复制标题

DOI:
10.1186/1471-2105-12-69
复制
发表时间:
2011-03-08
期刊:
影响因子:
3
通讯作者:
Stenius U
Stenius U
中科院分区:
生物学4区
文献类型:
--
作者:
Guo Y;Korhonen A;Liakata M;Silins I;Hogberg J;Stenius U

文献摘要

参考文献

被引文献

相似文献

生物医学中的许多实际任务需要访问科学文献中的特定类型的信息;例如,有关研究结果或结论的信息。已经开发了几种方案来描述科学期刊文章中的此类信息。例如,一个简单的基于章节的方案将摘要中的单个句子分配到目标、方法、结果和结论等章节下。文本信息结构的一些方案已被证明是有用的生物医学文本挖掘(BIO-TM)的任务(如自动摘要)。然而,以用户为中心的评价在现实生活中的任务一直缺乏。我们采取了三种不同类型和粒度的方案-基于节名,论证区(AZ)和核心科学概念(CoreSC)-并评估其实用性的现实生活中的任务,重点是生物医学摘要:癌症风险评估(CRA)。我们根据每个方案注释CRA摘要的语料库,开发分类器用于自动识别摘要中的方案,并直接在CRA的上下文中评估手动和自动分类。我们的研究结果表明,对于每个方案,大多数类别出现在摘要中,尽管其中两个方案(AZ和CoreSC)最初是为完整的期刊文章开发的。所有的方案都可以使用机器学习相对可靠地在摘要中识别。此外,当癌症风险评估员被呈现有方案注释的摘要时,他们发现相关信息的速度比呈现无注释的摘要时要快得多,即使注释是使用自动分类器产生的。有趣的是,在这个基于用户的评估中,基于节名的粗粒度方案被证明对CRA几乎和最细粒度的CoreSC方案一样有用。我们已经表明,现有的计划,旨在捕捉科学文献的信息结构,可以应用于生物医学摘要,并可以自动识别,在他们的准确性是足够高的,有利于在生物医学的现实生活中的任务。
Many practical tasks in biomedicine require accessing specific types of information in scientific literature; e.g. information about the results or conclusions of the study in question. Several schemes have been developed to characterize such information in scientific journal articles. For example, a simple section-based scheme assigns individual sentences in abstracts under sections such as Objective, Methods, Results and Conclusions. Some schemes of textual information structure have proved useful for biomedical text mining (BIO-TM) tasks (e.g. automatic summarization). However, user-centered evaluation in the context of real-life tasks has been lacking. We take three schemes of different type and granularity - those based on section names, Argumentative Zones (AZ) and Core Scientific Concepts (CoreSC) - and evaluate their usefulness for a real-life task which focuses on biomedical abstracts: Cancer Risk Assessment (CRA). We annotate a corpus of CRA abstracts according to each scheme, develop classifiers for automatic identification of the schemes in abstracts, and evaluate both the manual and automatic classifications directly as well as in the context of CRA. Our results show that for each scheme, the majority of categories appear in abstracts, although two of the schemes (AZ and CoreSC) were developed originally for full journal articles. All the schemes can be identified in abstracts relatively reliably using machine learning. Moreover, when cancer risk assessors are presented with scheme annotated abstracts, they find relevant information significantly faster than when presented with unannotated abstracts, even when the annotations are produced using an automatic classifier. Interestingly, in this user-based evaluation the coarse-grained scheme based on section names proved nearly as useful for CRA as the finest-grained CoreSC scheme. We have shown that existing schemes aimed at capturing information structure of scientific documents can be applied to biomedical abstracts and can be identified in them automatically with an accuracy which is high enough to benefit a real-life task in biomedicine.
DOI: 10.1186/1471-2105-10-303
发表时间: 2009-09-22
期刊: BMC bioinformatics
影响因子: 3
作者:
Korhonen A;Silins I;Sun L;Stenius U
通讯作者: Stenius U
DOI: 10.2307/2529310
发表时间: 1977-01-01
期刊: BIOMETRICS
影响因子: 1.9
作者:
LANDIS, JR;KOCH, GG
通讯作者: KOCH, GG
DOI: 10.1214/aoms/1177730491
发表时间: 1947-01-01
影响因子: --
作者:
MANN, HB;WHITNEY, DR
通讯作者: WHITNEY, DR
DOI: 10.1177/001316446002000104
发表时间: 1960-01-01
影响因子: 2.7
作者:
COHEN, J
通讯作者: COHEN, J
DOI: 10.1186/1471-2105-9-193
发表时间: 2008-04-14
期刊: BMC bioinformatics
影响因子: 3
作者:
Karamanis N;Seal R;Lewin I;McQuilton P;Vlachos A;Gasperin C;Drysdale R;Briscoe T
通讯作者: Briscoe T