Automatic extraction of keywords from scientific text: application to the knowledge domain of protein families

Automatic extraction of keywords from scientific text: application to the knowledge domain of protein families
复制标题

DOI:
10.1093/bioinformatics/14.7.600
复制
发表时间:
1998-01-01
期刊:
影响因子:
5.8
通讯作者:
Valencia, A
Valencia, A
中科院分区:
生物学3区
文献类型:
--
作者:
Andrade, MA;Valencia, A

文献摘要

被引文献

相似文献

动机:注释不同蛋白质序列的生物功能是目前由人类专家执行的一个耗时的过程。基因组分析工具在执行这项任务时遇到了很大的困难。数据库馆长、基因组分析工具的开发者和生物学家一般都可以从能够建议功能注释的工具的访问中受益,并促进对功能信息的访问。方法:我们在这里展示了蛋白质功能自动注释系统的第一个原型。该系统由与给定蛋白质相关的摘要集合触发,它能够直接从科学文献中提取生物信息,即MEDLINE摘要。通过与领域特定背景分布相比较的相对累积来选择相关关键字。同时,选择最具代表性的句子和MEDLINE摘要呈现给最终用户:进化信息被认为是蛋白质功能领域的主要特征。结果:本系统对不同的蛋白质家族进行了测试,并详细讨论了其中的三个例子:“共济失调-毛细血管扩张相关蛋白”、“RAN GTP酶”和“碳酸氢酶”。我们发现,提供给系统的信息量和注释的质量之间总体上有很好的相关性。最后,讨论了系统目前的局限性和未来的发展方向。可用性:目前的系统可以看作是一个原型系统。因此,可以在http://columba.ebi.ac.uk:8765/andrade/abx.上作为服务器访问它该系统接受与要评估的一个或多个蛋白质有关的测试(最好是按关键字进行MEDLINE搜索的结果),结果以网页的形式返回,以查找关键字、句子和摘要。补充信息:包含正文中提到的例子的箔信息的网页可在http://www.cnb.uam.es/similar到cnbprot/关键字/联系方式:valencia@cnb.uam.es获得。
Motivation: Annotation of the biological function of different protein sequences is a time-consuming process currently performed by human experts. Genome analysis tools encounter great difficulty in performing this task. Database curators, developers of genome analysis tools and biologists in general could benefit from access to tools able to suggest functional annotations and facilitate access to functional information.Approach: We present here the first prototype of a system for the automatic annotation of protein function. The system is triggered by collections of abstracts related to a given protein, and it is able to extract biological information directly from scientific literature, i.e. MEDLINE abstracts. Relevant keywords are selected by their relative accumulation in comparison with a domain-specific background distribution. Simultaneously the most representative sentences and MEDLINE abstracts are selected and presented to the end-user: Evolutionary information is considered as a predominant characteristic in the domain of protein function. Our system consequently extracts domain-specific information from the analysis of a set of protein families.Results: The system has been tested with differ-ent protein families, of which three examples are discussed in detail here: 'ataxia-telangiectasia associated protein', 'ran GTPase' and 'carbonic anhydrase'. We found generally good correlation between the amount of information provided to the system and the quality of the annotations. Finally, the current limitations and future developments of the system al-e discussed.Availability: The current system can be considered as a prototype system. As such, it can be accessed as a server at http://columba.ebi.ac.uk:8765/andrade/abx. The system accepts test related to the protein or proteins to be evaluated (optimally, the result of a MEDLINE search by keyword) and the results are returned in the form of Web pages for keywords, sentences and abstracts. Supplementary information: Web pages containing foil information on the examples mentioned in the text are available at: http://www.cnb.uam.es/similar to cnbprot/keywords/Contact: valencia@cnb.uam.es.