Information extraction from full text scientific articles: where are the keywords?

Information extraction from full text scientific articles: where are the keywords?
复制标题

DOI:
10.1186/1471-2105-4-20
复制
发表时间:
2003-05-29
期刊:
影响因子:
3
通讯作者:
Andrade MA
Andrade MA
中科院分区:
生物学4区
文献类型:
--
作者:
Shah PK;Perez-Iratxeta C;Bork P;Andrade MA

文献摘要

被引文献

相似文献

迄今为止,许多从科学文章中提取生物信息的方法仅限于文章的摘要。不过,目前已有电子版全文文章,提供了更大的数据来源。出现了一些问题,例如扫描全文文章的努力是否值得,或者从文章的不同部分提取的信息是否相关。在这项工作中,我们解决了这些问题,表明标准科学文章不同部分(摘要、引言、方法、结果和讨论)的关键词内容非常异构。尽管摘要包含关键词与总单词数的最佳比例,但文章的其他部分可能是生物学相关数据的更好来源。
To date, many of the methods for information extraction of biological information from scientific articles are restricted to the abstract of the article. However, full text articles in electronic version, which offer larger sources of data, are currently available. Several questions arise as to whether the effort of scanning full text articles is worthy, or whether the information that can be extracted from the different sections of an article can be relevant. In this work we addressed those questions showing that the keyword content of the different sections of a standard scientific article (abstract, introduction, methods, results, and discussion) is very heterogeneous. Although the abstract contains the best ratio of keywords per total of words, other sections of the article may be a better source of biologically relevant data.