What the papers say: text mining for genomics and systems biology.

What the papers say: text mining for genomics and systems biology.
复制标题

DOI:
10.1186/1479-7364-5-1-17
复制
发表时间:
2010-10
期刊:
影响因子:
4.5
通讯作者:
Stumpf MP
Stumpf MP
中科院分区:
医学3区
文献类型:
--
作者:
Harmston N;Filsell W;Stumpf MP

文献摘要

相似文献

对于大多数科学家来说,跟上快速增长的文献几乎是不可能的。这可能会产生可怕的后果。首先,我们可能会因为无法再可靠地掌握已发表的文献而将研究时间和资源浪费在重新发明轮子上。其次,也许更有害的是,来自不同科学学科的知识的明智(或偶然)组合需要遵循不同和不同的研究文献,即使对于研究出版物的最热心读者来说,也很快变得不可能。文本挖掘——从(电子)发布的来源中自动提取信息——可能会发挥重要作用——但前提是我们知道如何利用其优势并克服其弱点。由于我们预计科学成果的发表速度不会下降,因此文本挖掘工具现在变得至关重要,以应对这种信息爆炸并从中获得最大利益。在基因组学中,随着越来越多的罕见致病变异被发现并需要了解,这一点尤其紧迫。不熟悉这项技术可能会使科学家和生物医学监管者处于严重不利地位。在这篇综述中,我们介绍了现代文本挖掘的基本概念及其在基因组学和系统生物学中的应用。我们希望这次审查能够达到三个目的:(i)对该领域的现状提供及时且有用的概述,包括对当前挑战的调查; (ii) 使研究人员能够决定如何以及何时在自己的研究中应用文本挖掘工具; (iii) 强调基因组学和系统生物学研究界如何帮助使生物医学摘要和文本的文本挖掘变得更加简单。
Keeping up with the rapidly growing literature has become virtually impossible for most scientists. This can have dire consequences. First, we may waste research time and resources on reinventing the wheel simply because we can no longer maintain a reliable grasp on the published literature. Second, and perhaps more detrimental, judicious (or serendipitous) combination of knowledge from different scientific disciplines, which would require following disparate and distinct research literatures, is rapidly becoming impossible for even the most ardent readers of research publications. Text mining -- the automated extraction of information from (electronically) published sources -- could potentially fulfil an important role -- but only if we know how to harness its strengths and overcome its weaknesses. As we do not expect that the rate at which scientific results are published will decrease, text mining tools are now becoming essential in order to cope with, and derive maximum benefit from, this information explosion. In genomics, this is particularly pressing as more and more rare disease-causing variants are found and need to be understood. Not being conversant with this technology may put scientists and biomedical regulators at a severe disadvantage. In this review, we introduce the basic concepts underlying modern text mining and its applications in genomics and systems biology. We hope that this review will serve three purposes: (i) to provide a timely and useful overview of the current status of this field, including a survey of present challenges; (ii) to enable researchers to decide how and when to apply text mining tools in their own research; and (iii) to highlight how the research communities in genomics and systems biology can help to make text mining from biomedical abstracts and texts more straightforward.