Sentence retrieval for abstracts of randomized controlled trials

Sentence retrieval for abstracts of randomized controlled trials
复制标题

DOI:
10.1186/1472-6947-9-10
复制
发表时间:
2009-01-10
影响因子:
3.5
通讯作者:
Chung, Grace Y.
Chung, Grace Y.
中科院分区:
医学3区
文献类型:
--
作者:
Chung, Grace Y.

文献摘要

被引文献

相似文献

背景:循证医学(EBM)的实践要求临床医生将他们的专业知识与最新的科学研究相结合。但随着发表的文章数量不断增加,这变得越来越困难。显然需要更好的工具来提高临床医生搜索原始文献的能力。随机临床试验(RCT)是记录治疗方案有效性的最可靠的证据来源。本文描述了从随机对照试验摘要中检索关键句子,作为帮助用户找到有关临床研究实验设计的相关事实的一个步骤。方法:使用条件随机场(CRF)(一种流行且成功的自然语言处理问题方法),对涉及干预、参与者和结果测量的句子进行自动分类。这是通过扩展先前在摘要中标记句子的方法来完成的,该方法用于与科学论证或修辞角色相关的一般类别:目标、方法、结果和结论。方法在多个 RCT 摘要语料库上进行了测试。使用第一个结构化摘要,其标题专门表明干预、参与者和结果测量。此外,还准备了手动注释的结构化和非结构化摘要语料库,用于测试识别属于每个类别的句子的分类器。结果:使用条件随机场,可以将句子标记为四种修辞角色,F 分数为 0.93-0.98。这优于支持向量机的使用。此外,在非结构化和结构化摘要中,可以自动标记干预、参与者和结果测量的句子,其中章节标题没有具体指示这三个主题。干预和结果测量句子的 F 分数高达 0.83 和 0.84。 结论:结果表明,在结构化和非结构化摘要报告中,RCT 的一些方法论要素在句子级别上都是可识别的。这是有希望的,因为自动标记的句子可能会形成简洁的摘要,有助于信息检索和更细粒度的提取。
Background: The practice of evidence-based medicine (EBM) requires clinicians to integrate their expertise with the latest scientific research. But this is becoming increasingly difficult with the growing numbers of published articles. There is a clear need for better tools to improve clinician's ability to search the primary literature. Randomized clinical trials (RCTs) are the most reliable source of evidence documenting the efficacy of treatment options. This paper describes the retrieval of key sentences from abstracts of RCTs as a step towards helping users find relevant facts about the experimental design of clinical studies.Method: Using Conditional Random Fields (CRFs), a popular and successful method for natural language processing problems, sentences referring to Intervention, Participants and Outcome Measures are automatically categorized. This is done by extending a previous approach for labeling sentences in an abstract for general categories associated with scientific argumentation or rhetorical roles: Aim, Method, Results and Conclusion. Methods are tested on several corpora of RCT abstracts. First structured abstracts with headings specifically indicating Intervention, Participant and Outcome Measures are used. Also a manually annotated corpus of structured and unstructured abstracts is prepared for testing a classifier that identifies sentences belonging to each category.Results: Using CRFs, sentences can be labeled for the four rhetorical roles with F-scores from 0.93-0.98. This outperforms the use of Support Vector Machines. Furthermore, sentences can be automatically labeled for Intervention, Participant and Outcome Measures, in unstructured and structured abstracts where the section headings do not specifically indicate these three topics. F-scores of up to 0.83 and 0.84 are obtained for Intervention and Outcome Measure sentences.Conclusion: Results indicate that some of the methodological elements of RCTs are identifiable at the sentence level in both structured and unstructured abstract reports. This is promising in that sentences labeled automatically could potentially form concise summaries, assist in information retrieval and finer-grained extraction.