Automatic reconstruction of a bacterial regulatory network using Natural Language Processing

Automatic reconstruction of a bacterial regulatory network using Natural Language Processing
复制标题

DOI:
10.1186/1471-2105-8-293
复制
发表时间:
2007-08-07
期刊:
影响因子:
3
通讯作者:
Collado-Vides, Julio
Collado-Vides, Julio
中科院分区:
生物学4区
文献类型:
--
作者:
Rodriguez-Penagos, Carlos;Salgado, Heladia;Collado-Vides, Julio

文献摘要

被引文献

相似文献

背景:生物数据库的人工管理是一个昂贵且劳动密集型的过程,对于高质量的综合数据至关重要。在本文中,我们报告了一个最先进的自然语言处理系统的实现,该系统直接从不同的摘要和全文论文集合中创建计算机可读的监管交互网络。我们的主要目标是了解如何使用文本挖掘技术的自动注释可以补充人工管理的生物数据库。我们实现了一个基于规则的系统,从不同的文件集处理在大肠杆菌K-12的调控生成网络。结果:性能评估是基于最全面的转录调控数据库的任何生物体,手动策划的RegulonDB,其中45%,我们能够自动重建。从我们的自动分析中,我们还能够从尚未策展的论文中发现一些新的相互作用,或者在手动过滤和文献综述中遗漏的相互作用。我们还提出了一种新的监管交互标记语言,它比SBML更适合同时表示生物学家和文本挖掘者感兴趣的数据。对文本自动处理的输出进行手动管理是补充更详细的文献综述的好方法,无论是为了验证已经注释的结果,或者发现在分类或管理阶段可能被忽视的事实和信息。
Background: Manual curation of biological databases, an expensive and labor-intensive process, is essential for high quality integrated data. In this paper we report the implementation of a state-of-the-art Natural Language Processing system that creates computer-readable networks of regulatory interactions directly from different collections of abstracts and full-text papers. Our major aim is to understand how automatic annotation using Text-Mining techniques can complement manual curation of biological databases. We implemented a rule-based system to generate networks from different sets of documents dealing with regulation in Escherichia coli K-12.Results: Performance evaluation is based on the most comprehensive transcriptional regulation database for any organism, the manually-curated RegulonDB, 45% of which we were able to recreate automatically. From our automated analysis we were also able to find some new interactions from papers not already curated, or that were missed in the manual filtering and review of the literature. We also put forward a novel Regulatory Interaction Markup Language better suited than SBML for simultaneously representing data of interest for biologists and text miners.Conclusion: Manual curation of the output of automatic processing of text is a good way to complement a more detailed review of the literature, either for validating the results of what has been already annotated, or for discovering facts and information that might have been overlooked at the triage or curation stages.