SystemT: A System for Declarative Information Extraction

SystemT: A System for Declarative Information Extraction
复制标题

DOI:
10.1145/1519103.1519105
复制
发表时间:
2008-12-01
期刊:
影响因子:
1.1
通讯作者:
Zhu, Huaiyu
Zhu, Huaiyu
中科院分区:
计算机科学4区
文献类型:
--
作者:
Krishnamurthy, Rajasekar;Li, Yunyao;Zhu, Huaiyu

文献摘要

被引文献

相似文献

随着企业内外的应用程序遇到越来越多的非结构化数据,人们对信息提取(IE)领域重新产生了兴趣——这是一门从非结构化文本中提取结构化信息的学科。NLP社区开发的经典IE技术是基于级联语法和正则表达式的。然而,由于基于语法的提取的固有局限性,这些技术无法:(i)扩展到大型数据集,(ii)支持复杂信息任务的表达性要求。在IBM阿尔马登研究中心,我们正在开发SystemT,一个通过采用代数方法解决这些限制的IE系统。通过利用众所周知的数据库概念,如声明性查询和基于成本的优化,SystemT支持复杂信息提取任务的可伸缩执行。在本文中,我们提出了SystemT方法来进行信息提取。我们描述了我们的提取代数,并展示了我们的优化技术在复杂提取任务的运行时间上提供数量级减少的有效性。
As applications within and outside the enterprise encounter increasing volumes of unstructured data, there has been renewed interest in the area of information extraction (IE) - the discipline concerned with extracting structured information from unstructured text. Classical IE techniques developed by the NLP community were based on cascading grammars and regular expressions. However, due to the inherent limitations of grammar-based extraction, these techniques are unable to: (i) scale to large data sets, and (ii) support the expressivity requirements of complex information tasks. At the IBM Almaden Research Center, we are developing SystemT, an IE system that addresses these limitations by adopting an algebraic approach. By leveraging well-understood database concepts such as declarative queries and cost-based optimization, SystemT enables scalable execution of complex information extraction tasks. In this paper, we motivate the SystemT approach to information extraction. We describe our extraction algebra and demonstrate the effectiveness of our optimization techniques in providing orders of magnitude reduction in the running time of complex extraction tasks.