Automatically linking MEDLINE abstracts to the Gene Ontology

Automatically linking MEDLINE abstracts to the Gene Ontology
复制标题

DOI:
--
复制
发表时间:
2003
期刊:
--
影响因子:
--
通讯作者:
T. Smith;J. Cleary
T. Smith;J. Cleary
中科院分区:
其他
文献类型:
--
作者:
T. Smith;J. Cleary

文献摘要

被引文献

相似文献

最近已经写了很多关于需要有效的工具和方法来挖掘生物医学文献中存在的丰富信息的文章(Mack和Hehenberger,2002; Blagosklonny和Pardee,2001; Rindflesch等人,2002)-概念生物学的活动。在大型电子文档存储(如PubMed和PNAS)上运行的关键词搜索引擎提供了一些帮助,但存在限制其有效性的根本障碍。首先,科学家们在描述基因、蛋白质、药物、疾病、组织和疗法的研究时,对使用何种术语没有普遍的共识,这使得很难制定一个检索正确文档的搜索查询。其次,查找相关文章只是调查过程的一个方面。一个更基本的目标是在已发表文献中存在的事实之间建立联系和关系,以“验证当前的假设或产生新的假设”(巴恩斯和罗伯逊,2002年)-关键字搜索引擎几乎不支持。一个有前途的解决方案是将生物医学文献纳入基因本体论(GO)的结构化组织(Consortium,2000)。大量的基因组/蛋白质组数据库(例如SwissProt、SGD、InterPro、FlyBase等)以某种方式利用GO来链接和统一表达数据,将基因和蛋白质组织成或多或少一致的功能组,并解决命名中的一些模糊性,但在直接利用GO与文档方面几乎没有取得进展。例如,本文作者在2002年中期进行了大量的检索,在公共数据库中发现与GO术语直接或间接相关的MEDLINE摘要不到3万篇。在过去的一年里,情况有了很大的改善,最近的一次检索(2003年4月完成)发现了大约120,000篇与基因本体论相关的MEDLINE摘要,但如果继续手工操作,MEDLINE数据库1中包含的所有600万篇摘要都与GO术语相关联,还需要很长时间。
Much has been written recently about the need for effective tools and methods for mining the wealth of information present in biomedical literature (Mack and Hehenberger, 2002; Blagosklonny and Pardee, 2001; Rindflesch et al., 2002)—the activity of conceptual biology. Keyword search engines operating over large electronic document stores (such as PubMed and the PNAS) offer some help, but there are fundamental obstacles that limit their effectiveness. In the first instance, there is no general consensus among scientists about the vernacular to be used when describing research about genes, proteins, drugs, diseases, tissues and therapies, making it very difficult to formulate a search query that retrieves the right documents. Secondly, finding relevant articles is just one aspect of the investigative process. A more fundamental goal is to establish links and relationships between facts existing in published literature in order to “validate current hypotheses or to generate new ones” (Barnes and Robertson, 2002)—something keyword search engines do little to support. One promising solution is to bring biomedical literature into the structured organisation of the Gene Ontology (GO) (Consortium, 2000). A large number of genomic/proteomic databases (e.g. SwissProt, SGD, InterPro, FlyBase, etc) make use of GO in some way to link and unify expression data, organize genes and proteins into more or less coherent functional groups, and resolve some of the ambiguities in nomenclature, but little progress has been made towards exploiting GO directly with documents. For example, a substantial search effort made by the authors of this paper in mid-2002 found fewer than thirty thousand MEDLINE abstracts directly or indirectly linked to GO terms in public databases. The situation has improved greatly over the past year, such that a more recent search (completed in April 2003) uncovered about 120,000 MEDLINE abstracts linked to the Gene Ontology, but it will still take a very long time before all six million abstracts contained in the MEDLINE database1 are associated with GO terms if the process continues to be done manually.