Unsupervised named-entity extraction from the Web: An experimental study

Unsupervised named-entity extraction from the Web: An experimental study
复制标题

DOI:
10.1016/j.artint.2005.03.001
复制
发表时间:
2005-06-01
影响因子:
14.4
通讯作者:
Yates, A
Yates, A
中科院分区:
计算机科学2区
文献类型:
--
作者:
Etzioni, O;Cafarella, M;Yates, A

文献摘要

被引文献

相似文献

KNOWITALL 系统旨在以无监督、独立于领域和可扩展的方式自动化从网络中提取大量事实(例如科学家或政治家的姓名)的繁琐过程。本文概述了 KNOWITALL 的新颖架构和设计原则,强调其无需任何手工标记的训练示例即可提取信息的独特能力。在第一次主要运行中,KNOWITALL 提取了超过 50,000 个类实例,但提出了一个挑战:我们如何在不牺牲精度的情况下提高 KNOWITALL 的召回率和提取率?本文提出了三种不同的方法来应对这一挑战并评估其性能。模式学习学习特定领域的提取规则,从而实现额外的提取。子类提取自动识别子类以提高召回率(例如,“化学家”和“生物学家”被识别为“科学家”的子类)。列表提取定位类实例列表,学习每个列表的“包装器”,并提取每个列表的元素。由于每种方法都是从 KNOWITALL 的与领域无关的方法引导,因此这些方法也避免了手工标记的训练示例。该论文报告了实验,重点关注建立命名实体列表,衡量每种方法的相对功效并证明其协同作用,我们的方法使 KNOWITALL 的召回率提高了 4 倍到 8 倍,精度为 0.90,并发现了 Tipster 地名词典中缺失的 10,000 多个城市。
The KNOWITALL system aims to automate the tedious process of extracting large collections of facts (e.g., names of scientists or politicians) from the Web in an unsupervised, domain-independent, and scalable manner. The paper presents an overview of KNOWITALL's novel architecture and design principles, emphasizing its distinctive ability to extract information without any hand-labeled training examples. In its first major run, KNOWITALL extracted over 50,000 class instances, but suggested a challenge: How can we improve KNOWITALL's recall and extraction rate without sacrificing precision? This paper presents three distinct ways to address this challenge and evaluates their performance. Pattern Learning learns domain-specific extraction rules, which enable additional extractions. Subclass Extraction automatically identifies sub-classes in order to boost recall (e.g., "chemist" and c biologist" are identified as sub-classes of "scientist"). List Extraction locates lists of class instances, learns a "wrapper" for each list, and. extracts elements of each list. Since each method bootstraps from KNOWITALL's domain-independent methods, the methods also obviate hand-labeled training examples. The paper reports on experiments, focused on building lists of named entities, that measure the relative efficacy of each method and demonstrate their synergy. In concert, our methods gave KNOWITALL a 4-fold to 8-fold increase in recall at precision of 0.90, and discovered over 10,000 cities missing from the Tipster Gazetteer.