An automated method to build a corpus of rhetorically-classified sentences in biomedical texts

An automated method to build a corpus of rhetorically-classified sentences in biomedical texts
复制标题

一种在生物医学文本中构建修辞分类句子语料库的自动化方法

DOI:
10.3115/v1/w14-2103
复制
发表时间:
2014
期刊:
AMIA ... Annual Symposium proceedings. AMIA Symposium
影响因子:
--
通讯作者:
Robert E. Mercer
Robert E. Mercer
中科院分区:
--
文献类型:
--
作者:
Hospice Houngbo;Robert E. Mercer

文献摘要

被引文献

相似文献

生物医学文本中句子的修辞分类是科学论证成分识别的一项重要任务。生成有监督的机器学习模型来进行这种识别需要对修辞类别介绍(或背景)、方法、结果、讨论(或结论)进行标注的语料库。目前,存在一些小型的带注释的语料库。我们使用“This”一词的共指文本的直接特征来构建一个从大型生物医学研究论文数据集中提取的自标注语料库。除了引言外,语料库对所有修辞类别都进行了标注,而不涉及领域专家。在10次交叉验证中,我们报告总体F
The rhetorical classification of sentences in biomedical texts is an important task in the recognition of the components of a scientific argument. Generating supervised machine learned models to do this recognition requires corpora annotated for the rhetorical categories Introduction (or Background), Method, Result, Discussion (or Conclusion). Currently, a few, small annotated corpora exist. We use a straightforward feature of co-referring text using the word “this” to build a selfannotating corpus extracted from a large biomedical research paper dataset. The corpus is annotated for all of the rhetorical categories except Introduction without involving domain experts. In a 10-fold cross-validation, we report an overall F