Recognizing speculative language in biomedical research articles: a linguistically motivated perspective.

Recognizing speculative language in biomedical research articles: a linguistically motivated perspective.
复制标题

DOI:
10.1186/1471-2105-9-s11-s10
复制
发表时间:
2008-11-19
期刊:
影响因子:
3
通讯作者:
Bergler S
Bergler S
中科院分区:
生物学4区
文献类型:
--
作者:
Kilicoglu H;Bergler S

文献摘要

被引文献

相似文献

由于科学方法论的性质,研究文章中充满了推测性和尝试性的陈述,也被称为模糊限制语。我们探索一种基于语言学动机的方法来解决生物医学研究文章中识别此类语言的问题。我们的方法借鉴了以前的语言学工作,以及现有的词汇资源,以创建一个字典的模糊限制语的线索,并通过引入句法模式来扩展它。此外,认识到对冲线索不同的投机强度,我们分配他们的权重有两种方式:自动使用信息增益(IG)的措施和半自动的基础上,他们的类型和中心对冲。模糊限制语线索的权重被用来确定句子的推测强度。我们在两个公开的对冲数据集上测试我们的系统。在果蝇数据集上,我们使用半自动加权方案实现了0.85的精确度-召回率盈亏平衡点(BEP),使用信息增益加权方案实现了0.80的较低BEP。这些结果与先前报告的最佳结果(BEP为0.85)具有竞争力。在BMC数据集上,使用半自动加权产生的BEP为0.82,与先前报告的最佳结果(BEP为0.76)相比,有统计学显著改善(p <0.01),而信息增益加权产生的BEP为0.70。我们的研究结果表明,投机性语言可以成功地识别与语言动机的方法,并证实,选择的对冲工具影响句子的投机强度,这可以通过加权对冲线索合理地捕获。在BMC数据集上使用半自动加权方案获得的改进表明,我们的面向语言的方法比基于机器学习的方法更具可移植性。用信息增益加权方案获得的较低性能表明,该方法可以受益于用于自动诱导权重的较大的手动注释语料库。
Due to the nature of scientific methodology, research articles are rich in speculative and tentative statements, also known as hedges. We explore a linguistically motivated approach to the problem of recognizing such language in biomedical research articles. Our approach draws on prior linguistic work as well as existing lexical resources to create a dictionary of hedging cues and extends it by introducing syntactic patterns. Furthermore, recognizing that hedging cues differ in speculative strength, we assign them weights in two ways: automatically using the information gain (IG) measure and semi-automatically based on their types and centrality to hedging. Weights of hedging cues are used to determine the speculative strength of sentences. We test our system on two publicly available hedging datasets. On the fruit-fly dataset, we achieve a precision-recall breakeven point (BEP) of 0.85 using the semi-automatic weighting scheme and a lower BEP of 0.80 with the information gain weighting scheme. These results are competitive with the previously reported best results (BEP of 0.85). On the BMC dataset, using semi-automatic weighting yields a BEP of 0.82, a statistically significant improvement (p <0.01) over the previously reported best result (BEP of 0.76), while information gain weighting yields a BEP of 0.70. Our results demonstrate that speculative language can be recognized successfully with a linguistically motivated approach and confirms that selection of hedging devices affects the speculative strength of the sentence, which can be captured reasonably by weighting the hedging cues. The improvement obtained on the BMC dataset with a semi-automatic weighting scheme indicates that our linguistically oriented approach is more portable than the machine-learning based approaches. Lower performance obtained with the information gain weighting scheme suggests that this method may benefit from a larger, manually annotated corpus for automatically inducing the weights.