Beyond accuracy: creating interoperable and scalable text-mining web services

Beyond accuracy: creating interoperable and scalable text-mining web services
复制标题

DOI:
10.1093/bioinformatics/btv760
复制
发表时间:
2016-06-15
期刊:
影响因子:
5.8
通讯作者:
Lu, Zhiyong
Lu, Zhiyong
中科院分区:
生物学3区
文献类型:
--
作者:
Wei, Chih-Hsuan;Leaman, Robert;Lu, Zhiyong

文献摘要

被引文献

相似文献

摘要:生物医学文献是知识丰富的资源,也是未来研究的重要基础。 PubMed 上有超过 2400 万篇文章并且增长率不断上升,自动文本处理的研究变得越来越重要。我们在此报告我们最近开发的基于网络的文本挖掘服务,用于生物医学概念识别和标准化。与大多数文本挖掘软件工具不同,我们的网络服务集成了多种最先进的实体标记系统(DNorm、GNormPlus、SR4GN、tmChem 和 tmVar),并提供批处理模式,能够处理多种格式(例如 BioC)的任意文本输入(例如学术出版物、专利和医疗记录)。我们支持多种标准,使我们的服务具有互操作性,并允许与其他文本处理管道更简单地集成。为了最大限度地提高可扩展性,我们对所有 PubMed 文章进行了预处理,并使用计算机集群来处理任意文本的大量请求。
A Summary: The biomedical literature is a knowledge-rich resource and an important foundation for future research. With over 24 million articles in PubMed and an increasing growth rate, research in automated text processing is becoming increasingly important. We report here our recently developed web-based text mining services for biomedical concept recognition and normalization. Unlike most text-mining software tools, our web services integrate several state-of-the-art entity tagging systems (DNorm, GNormPlus, SR4GN, tmChem and tmVar) and offer a batch-processing mode able to process arbitrary text input (e.g. scholarly publications, patents and medical records) in multiple formats (e.g. BioC). We support multiple standards to make our service interoperable and allow simpler integration with other text-processing pipelines. To maximize scalability, we have preprocessed all PubMed articles, and use a computer cluster for processing large requests of arbitrary text.