The structural and content aspects of abstracts versus bodies of full text journal articles are different.

The structural and content aspects of abstracts versus bodies of full text journal articles are different.
复制标题

DOI:
10.1186/1471-2105-11-492
复制
发表时间:
2010-09-29
期刊:
影响因子:
3
通讯作者:
Hunter LE
Hunter LE
中科院分区:
生物学4区
文献类型:
--
作者:
Cohen KB;Johnson HL;Verspoor K;Roeder C;Hunter LE

文献摘要

参考文献

被引文献

相似文献

期刊文章全文工作的增加和 PubMedCentral 的发展有机会在生物医学文本挖掘的完成方式上产生重大范式转变。然而,到目前为止,还没有全面描述全文期刊文章的正文与摘要的区别,而摘要迄今为止一直是大多数生物医学文本挖掘研究的主题。我们检查了全文文章的摘要和正文的结构和语言方面、文本挖掘工具在两者上的性能,以及它们之间命名实体的各种语义类别的分布。我们发现了明显的结构差异,文章正文中的句子更长,正文中括号内材料的使用比摘要中更多。我们发现了语言特征方面的内容差异。我们检查的四分之三的语言特征在两种流派之间的分布在统计上存在显着差异。我们还发现了语义特征分布方面的内容差异。四分之三的语义类别的每千词密度存在显着差异,并且它们在两种类型中出现的程度也存在明显差异。关于文本挖掘工具的性能,我们发现突变查找器在两种类型中的表现同样出色,但各种基因提及系统在文章正文上的表现比在摘要上的表现要差得多。摘要中的词性标记也比文章正文中的词性标记更准确。文章摘要和文章正文在结构和内容方面存在显着差异。随着文本挖掘领域更多地进入处理全文文章的领域,其中许多差异可能会带来问题。然而,这些差异也为提取数据类型提供了许多机会,特别是在括号文本中发现的数据类型,这些数据类型存在于文章正文中,但不存在于文章摘要中。
An increase in work on the full text of journal articles and the growth of PubMedCentral have the opportunity to create a major paradigm shift in how biomedical text mining is done. However, until now there has been no comprehensive characterization of how the bodies of full text journal articles differ from the abstracts that until now have been the subject of most biomedical text mining research. We examined the structural and linguistic aspects of abstracts and bodies of full text articles, the performance of text mining tools on both, and the distribution of a variety of semantic classes of named entities between them. We found marked structural differences, with longer sentences in the article bodies and much heavier use of parenthesized material in the bodies than in the abstracts. We found content differences with respect to linguistic features. Three out of four of the linguistic features that we examined were statistically significantly differently distributed between the two genres. We also found content differences with respect to the distribution of semantic features. There were significantly different densities per thousand words for three out of four semantic classes, and clear differences in the extent to which they appeared in the two genres. With respect to the performance of text mining tools, we found that a mutation finder performed equally well in both genres, but that a wide variety of gene mention systems performed much worse on article bodies than they did on abstracts. POS tagging was also more accurate in abstracts than in article bodies. Aspects of structure and content differ markedly between article abstracts and article bodies. A number of these differences may pose problems as the text mining field moves more into the area of processing full-text articles. However, these differences also present a number of opportunities for the extraction of data types, particularly that found in parenthesized text, that is present in article bodies but not in article abstracts.
DOI: 10.1186/1471-2105-4-20
发表时间: 2003-05-29
期刊: BMC bioinformatics
影响因子: 3
作者:
Shah PK;Perez-Iratxeta C;Bork P;Andrade MA
通讯作者: Andrade MA
DOI: 10.1002/cfg.91
发表时间: 2001
影响因子: --
作者:
Blaschke, C;Valencia, A
通讯作者: Valencia, A
DOI: 10.1093/bioinformatics/bth386
发表时间: 2004-11-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Corney, DPA;Buxton, BF;Jones, DT
通讯作者: Jones, DT
生物公约概述:生物学信息提取的批判性评估。
DOI: 10.1186/1471-2105-6-s1-s1
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Hirschman L;Yeh A;Blaschke C;Valencia A
通讯作者: Valencia A
评估生物学的文本挖掘系统:第二次生物综合社区挑战的概述。
DOI: 10.1186/gb-2008-9-s2-s1
发表时间: 2008
期刊: Genome biology
影响因子: 12.3
作者:
Krallinger M;Morgan A;Smith L;Leitner F;Tanabe L;Wilbur J;Hirschman L;Valencia A
通讯作者: Valencia A