Development and Analysis of NLP Pipelines in Argo
Development and Analysis of NLP Pipelines in Argo
复制标题
Argo 中 NLP 管道的开发与分析
DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
S. Ananiadou
中科院分区:
文献类型:
--
作者:
Rafal Rak;Andrew Rowley;Jacob Carter;S. Ananiadou
Developing sophisticated NLP pipelines composed of multiple processing tools and components available through different providers may pose a challenge in terms of their interoperability. The Unstructured Information Management Architecture (UIMA) is an industry standard whose aim is to ensure such interoperability by defining common data structures and interfaces. The architecture has been gaining attention from industry and academia alike, resulting in a large volume of UIMA-compliant processing components. In this paper, we demonstrate Argo, a Web-based workbench for the development and processing of NLP pipelines/workflows. The workbench is based upon UIMA, and thus has the potential of using many of the existing UIMA resources. We present features, and show examples, of facilitating the distributed development of components and the analysis of processing results. The latter includes annotation visualisers and editors, as well as serialisation to RDF format, which enables flexible querying in addition to data manipulation thanks to the semantic query language SPARQL. The distributed development feature allows users to seamlessly connect their tools to workflows running in Argo, and thus take advantage of both the available library of components (without the need of installing them locally) and the analytical tools.