MeSH: a window into full text for document summarization.

MeSH: a window into full text for document summarization.
复制标题

DOI:
10.1093/bioinformatics/btr223
复制
发表时间:
2011-07-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Srinivasan P
Srinivasan P
中科院分区:
其他
文献类型:
--
作者:
Bhattacharya S;Ha-Thuc V;Srinivasan P

文献摘要

参考文献

被引文献

相似文献

动机:以前在生物医学文本挖掘领域的研究历来仅限于MEDLINE记录中的标题、摘要和元数据。最近的研究活动,如TREC基因组学和BioCreAtIVE,有力地指出了超越摘要而进入全文领域的优点。然而,全文的处理不仅在所需资源方面更昂贵,而且在准确性方面也更昂贵。由于全文包含阐述、上下文、对比、补充等修饰,因此存在更大的误报风险。受此启发,我们探索了一种在摘要和全文的极端之间提供折衷的方法。具体地说,我们创建只包含重要部分的全文文档的精简版本。从长远来看,我们的目标是探索将此类摘要用于文档检索和信息提取等功能。在这里,我们重点设计总结策略。特别是,我们探索了由训练有素的注释员手动分配给文档的网格术语的使用,作为从全文文档中选择重要文本片段的线索。结果:我们的实验证实了我们的方法能够挑选出重要的文本部分。使用Rouge方法进行评估,对于我们的基于网格术语的方法,我们能够获得最高的Rouge-1、Rouge-2和Rouge-Su4 F-分数分别为0.4150、0.1435和0.1782,而最高基线分数分别为0.3815、0.1353和0.1428。使用基于网格轮廓的策略,我们能够分别获得0.4320、0.1497和0.1887的最大腮红F分数。对基线和我们提出的策略的人工评估进一步证实了我们的方法从全文中选择重要句子的能力。联系人:sanmitra-bhattacharya@uiowa.edu;
Motivation: Previous research in the biomedical text-mining domain has historically been limited to titles, abstracts and metadata available in MEDLINE records. Recent research initiatives such as TREC Genomics and BioCreAtIvE strongly point to the merits of moving beyond abstracts and into the realm of full texts. Full texts are, however, more expensive to process not only in terms of resources needed but also in terms of accuracy. Since full texts contain embellishments that elaborate, contextualize, contrast, supplement, etc., there is greater risk for false positives. Motivated by this, we explore an approach that offers a compromise between the extremes of abstracts and full texts. Specifically, we create reduced versions of full text documents that contain only important portions. In the long-term, our goal is to explore the use of such summaries for functions such as document retrieval and information extraction. Here, we focus on designing summarization strategies. In particular, we explore the use of MeSH terms, manually assigned to documents by trained annotators, as clues to select important text segments from the full text documents. Results: Our experiments confirm the ability of our approach to pick the important text portions. Using the ROUGE measures for evaluation, we were able to achieve maximum ROUGE-1, ROUGE-2 and ROUGE-SU4 F-scores of 0.4150, 0.1435 and 0.1782, respectively, for our MeSH term-based method versus the maximum baseline scores of 0.3815, 0.1353 and 0.1428, respectively. Using a MeSH profile-based strategy, we were able to achieve maximum ROUGE F-scores of 0.4320, 0.1497 and 0.1887, respectively. Human evaluation of the baselines and our proposed strategies further corroborates the ability of our method to select important sentences from the full texts. Contact: sanmitra-bhattacharya@uiowa.edu; padmini-srinivasan@uiowa.edu
DOI: 10.1016/0306-4573(95)00052-i
发表时间: 1995-09-01
影响因子: 8.6
作者:
BRANDOW, R;MITZE, K;RAU, LF
通讯作者: RAU, LF
DOI: 10.1016/j.ipm.2007.01.018
发表时间: 2007-11-01
影响因子: 8.6
作者:
Ling, Xu;Jiang, Jing;Schatz, Bruce
通讯作者: Schatz, Bruce
DOI: 10.1504/ijdmb.2007.012967
发表时间: 2007-01-01
影响因子: 0.3
作者:
Reeve, Lawrence H.;Han, Hyoil
通讯作者: Han, Hyoil
DOI: 10.1186/1471-2105-7-220
发表时间: 2006-04-21
期刊: BMC bioinformatics
影响因子: 3
作者:
Sehgal AK;Srinivasan P
通讯作者: Srinivasan P
DOI: 10.1016/j.ipm.2003.10.006
发表时间: 2004-11-01
影响因子: 8.6
作者:
Radev, DR;Jing, HY;Tam, D
通讯作者: Tam, D