Development and Analysis of NLP Pipelines in Argo

Development and Analysis of NLP Pipelines in Argo
复制标题

Argo 中 NLP 管道的开发与分析

DOI:
--
复制
发表时间:
2013
期刊:
Annual Meeting of the Association for Computational Linguistics
影响因子:
--
通讯作者:
S. Ananiadou
S. Ananiadou
中科院分区:
--
文献类型:
--
作者:
Rafal Rak;Andrew Rowley;Jacob Carter;S. Ananiadou

文献摘要

被引文献

相似文献

开发由不同提供商提供的多个处理工具和组件组成的复杂NLP管道可能会在互操作性方面带来挑战。非结构化信息管理体系结构(UIMA)是一个行业标准,其目的是通过定义公共数据结构和接口来确保这种互操作性。该架构已经引起了业界和学术界的关注,从而产生了大量的UIMA兼容处理组件。在本文中,我们展示了Argo,一个基于Web的工作台,用于开发和处理NLP管道/工作流。该工作台基于UIMA,因此有可能使用许多现有的UIMA资源。我们提出的功能,并显示的例子,促进组件的分布式开发和处理结果的分析。后者包括注释可视化器和编辑器,以及序列化为RDF格式,这使得灵活的查询除了数据操作感谢语义查询语言SPARQL。分布式开发功能允许用户将他们的工具无缝连接到Argo中运行的工作流程,从而利用可用的组件库(无需在本地安装)和分析工具。
Developing sophisticated NLP pipelines composed of multiple processing tools and components available through different providers may pose a challenge in terms of their interoperability. The Unstructured Information Management Architecture (UIMA) is an industry standard whose aim is to ensure such interoperability by defining common data structures and interfaces. The architecture has been gaining attention from industry and academia alike, resulting in a large volume of UIMA-compliant processing components. In this paper, we demonstrate Argo, a Web-based workbench for the development and processing of NLP pipelines/workflows. The workbench is based upon UIMA, and thus has the potential of using many of the existing UIMA resources. We present features, and show examples, of facilitating the distributed development of components and the analysis of processing results. The latter includes annotation visualisers and editors, as well as serialisation to RDF format, which enables flexible querying in addition to data manipulation thanks to the semantic query language SPARQL. The distributed development feature allows users to seamlessly connect their tools to workflows running in Argo, and thus take advantage of both the available library of components (without the need of installing them locally) and the analytical tools.