Generating gene summaries from biomedical literature: A study of semi-structured summarization

Generating gene summaries from biomedical literature: A study of semi-structured summarization
复制标题

DOI:
10.1016/j.ipm.2007.01.018
复制
发表时间:
2007-11-01
影响因子:
8.6
通讯作者:
Schatz, Bruce
Schatz, Bruce
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ling, Xu;Jiang, Jing;Schatz, Bruce

文献摘要

被引文献

相似文献

通过基因组学和相关生物医学学科的科学发现积累的大部分知识都隐藏在海量的生物医学文献中。由于了解基因调控是生物医学研究的基础,根据文献总结关于基因的所有现有知识是非常必要的,以帮助生物学家消化这些文献。在本文中,我们提出了一种从生物医学文献中自动生成基因摘要的方法的研究。与大多数现有的自动文本摘要工作不同的是,生成的摘要通常是提取的句子列表,我们建议生成一个由涵盖基因特定语义方面的句子组成的半结构化摘要。这种半结构化的摘要更适合描述基因,并对自动文本摘要提出了特殊的挑战。我们提出了一种两阶段的方法来为给定的基因生成这样的摘要-首先检索关于基因的文章,然后为每个特定的语义方面提取句子。我们在第一阶段解决了基因名称变异的问题,并在第二阶段提出了几种不同的句子提取方法。我们使用包含20个基因的测试集对所提出的方法进行了评估。实验结果表明,该方法能够从生物医学文献中自动生成有用的半结构化基因摘要,并且优于通用的摘要方法。在所有已提出的句子提取方法中,对基因上下文进行建模的概率语言建模方法表现最好。(C)2007爱思唯尔有限公司。保留所有权利。
Most knowledge accumulated through scientific discoveries in genomics and related biomedical disciplines is buried in the vast amount of biomedical literature. Since understanding gene regulations is fundamental to biomedical research, summarizing all the existing knowledge about a gene based on literature is highly desirable to help biologists digest the literature. In this paper, we present a study of methods for automatically generating gene summaries from biomedical literature. Unlike most existing work on automatic text summarization, in which the generated summary is often a list of extracted sentences, we propose to generate a semi-structured summary which consists of sentences covering specific semantic aspects of a gene. Such a semi-structured summary is more appropriate for describing genes and poses special challenges for automatic text summarization. We propose a two-stage approach to generate such a summary for a given gene - first retrieving articles about a gene and then extracting sentences for each specified semantic aspect. We address the issue of gene name variation in the first stage and propose several different methods for sentence extraction in the second stage. We evaluate the proposed methods using a test set with 20 genes. Experiment results show that the proposed methods can generate useful semi-structured gene summaries automatically from biomedical literature, and our proposed methods outperform general purpose summarization methods. Among all the proposed methods for sentence extraction, a probabilistic language modeling approach that models gene context performs the best. (C) 2007 Elsevier Ltd. All rights reserved.