PNAD-CSS: a workbench for constructing a protein name abbreviation dictionary

PNAD-CSS: a workbench for constructing a protein name abbreviation dictionary
复制标题

DOI:
10.1093/bioinformatics/16.2.169
复制
发表时间:
2000-02-01
期刊:
影响因子:
5.8
通讯作者:
Takagi, T
Takagi, T
中科院分区:
生物学3区
文献类型:
--
作者:
Yoshida, M;Fukuda, K;Takagi, T

文献摘要

被引文献

相似文献

动机:自最初开发以来,分子水平数据的数据库的整合和构建取得了进展。虽然生物分子彼此相关并形成一个复杂的系统,但信息存储在大量的文献档案或各种数据库中。生物学对象没有统一的命名规范,生物学术语可能具有歧义或多义性。这使得数据库的集成和交互变得困难。为了消除这些问题,机器可读的自然语言资源似乎是相当有前途的。结果:我们开发了蛋白质名称缩写词典构建支持系统(PNAD-CSS),该系统提供了各种便利的工具,可以降低构建蛋白质名称缩写词典的成本,该词典的条目来自生物医学论文的摘要。该系统使用户能够集中精力于更高层次的解释,消除了一些麻烦的任务,如管理摘要,提取蛋白质名称及其缩写,等等。为了提取一对蛋白质名称和缩写,我们已经开发了一个混合系统组成的PROPER系统和PNAD系统。PNAD系统可以从蛋白质名称中的括号-释义中提取蛋白质对,PROPER系统识别出这些蛋白质对,准确率为98.95%,召回率为95.56%,完全准确率为97.58%。http://www.hgc.ims.u-tokyo.ac.jp/service/tooldoc/KeX/intro.html其他软件也可应要求提供。联系作者:mikio@ims.u-tokyo.ac.jp。
Motivation: Since their initial development, integration and construction of databases for molecular-level data have progressed. Though biological molecules ave related to each other and form a complex system, the information is stored in the vast archives of the literature or in diverse databases. There is no unified naming convention for biological object, and biological terms may be ambiguous or polysemic. This makes the integration and interaction of databases difficult. In order to eliminate these problems, machine-readable natural language resources appear to be quite promising. We have developed a workbench for protein name abbreviation dictionary (PNAD) building.Results: We have developed PNAD Construction Support System (PNAD-CSS), which offers various convenient facilities to decrease the construction costs of a protein name abbreviation dictionary of which entries are collected from abstracts in biomedical papers. The system allows the users to concentrate on higher level interpretation by removing some troublesome tasks, e.g. management of abstracts, extracting protein names and their abbreviations, and so on. To extract a pair of protein names and abbreviations, we have developed a hybrid system composed of the PROPER System and the PNAD System. The PNAD System can extract the pairs from parenthetical-paraphrases involved in protein names, the PROPER System identified these pairs, with 98.95% precision, 95.56% recall and 97.58% complete precision.Availability: PROPER System is freely available from http://www.hgc.ims.u-tokyo.ac.jp/service/tooldoc/KeX/intro.html. The other software are also available on request. Contact the authors.Contact: mikio@ims.u-tokyo.ac.jp.