An automated method to build a corpus of rhetorically-classified sentences in biomedical texts
An automated method to build a corpus of rhetorically-classified sentences in biomedical texts
复制标题
一种在生物医学文本中构建修辞分类句子语料库的自动化方法
DOI:
10.3115/v1/w14-2103
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
Robert E. Mercer
中科院分区:
文献类型:
--
作者:
Hospice Houngbo;Robert E. Mercer
The rhetorical classification of sentences in biomedical texts is an important task in the recognition of the components of a scientific argument. Generating supervised machine learned models to do this recognition requires corpora annotated for the rhetorical categories Introduction (or Background), Method, Result, Discussion (or Conclusion). Currently, a few, small annotated corpora exist. We use a straightforward feature of co-referring text using the word “this” to build a selfannotating corpus extracted from a large biomedical research paper dataset. The corpus is annotated for all of the rhetorical categories except Introduction without involving domain experts. In a 10-fold cross-validation, we report an overall F