Text processing through web services: calling Whatizit

Text processing through web services: calling Whatizit
复制标题

DOI:
10.1093/bioinformatics/btm557
复制
发表时间:
2008-01-15
期刊:
影响因子:
5.8
通讯作者:
Jimeno, Antonio
Jimeno, Antonio
中科院分区:
生物学3区
文献类型:
--
作者:
Rebholz-Schuhmann, Dietrich;Arregui, Miguel;Jimeno, Antonio

文献摘要

被引文献

相似文献

动机:文本挖掘(TM)解决方案正在为生物医学研究界的研究人员提供有效的服务。此类解决方案必须随着资源数量和规模的增长(例如可用的受控词汇)的规模扩展,并且要处理大量文献(例如,PubMed中约有1700万个文档)以及用户社区的需求(例如,用于不同的方法。事实提取)。这些要求促使开发基于服务器的文献分析解决方案。Whatizit是一套模块,可以分析文本以获取包含的信息,例如任何科学出版物或MEDLINE摘要。特殊模块识别术语,然后将其链接到生物信息学数据库中的相应条目,例如UniprotkB/Swiss-Prot数据条目和基因本体论概念。其他模块确定了一组选定的注释类型,例如蛋白质ebimed分析管道产生的集合。对于MEDLINE摘要,Whatizit通过PMID或术语查询可以访问EBIS内部安装。对于大量用户自己的文本,可以在流媒体模式下操作服务器(http://www.ebi.ac.uk/webservices/whatizit)。
Motivation: Text-mining (TM) solutions are developing into efficient services to researchers in the biomedical research community. Such solutions have to scale with the growing number and size of resources (e.g. available controlled vocabularies), with the amount of literature to be processed (e.g. about 17 million documents in PubMed) and with the demands of the user community (e.g. different methods for fact extraction). These demands motivated the development of a server-based solution for literature analysis.Whatizit is a suite of modules that analyse text for contained information, e.g. any scientific publication or Medline abstracts. Special modules identify terms and then link them to the corresponding entries in bioinformatics databases such as UniProtKb/Swiss-Prot data entries and gene ontology concepts. Other modules identify a set of selected annotation types like the set produced by the EBIMed analysis pipeline for proteins. In the case of Medline abstracts, Whatizit offers access to EBIs in-house installation via PMID or term query. For large quantities of the users own text, the server can be operated in a streaming mode (http:// www.ebi.ac.uk/webservices/whatizit).