Part-of-Speech Tagging Using Progol

Part-of-Speech Tagging Using Progol
复制标题

使用 Progol 进行词性标注

DOI:
10.1007/3540635149_38
复制
发表时间:
1997
期刊:
International Conference on Inductive Logic Programming
影响因子:
--
通讯作者:
J. Cussens
J. Cussens
中科院分区:
--
文献类型:
--
作者:
J. Cussens

文献摘要

被引文献

相似文献

构建了一个用词性(POS)标签来“标记”单词的系统。该系统有两个组件:包含给定单词的一组可能的 POS 标签的词典,以及使用单词的上下文来消除单词的可能标签的规则。归纳逻辑编程 (ILP) 系统 Progol 用于以明确子句的形式归纳这些规则。最终理论包含 885 个条款。对于背景知识,Progol 使用简单的语法,其中标签是终结符,名词短语等谓词是非终结符。 Progol 被修改为允许缓存有关归纳过程中生成的子句的信息,这大大提高了效率。该系统对从不带引号的句子中提取的已知单词的每个单词的准确率达到了 96.4%。这与从相同数据得出的其他标记系统相当 [5,2,4],其准确度均在 96-97% 范围内。每句话准确率为 4 49.5%。
A system for ‘tagging’ words with their part-of-speech (POS) tags is constructed. The system has two components: a lexicon containing the set of possible POS tags for a given word, and rules which use a word's context to eliminate possible tags for a word. The Inductive Logic Programming (ILP) system Progol is used to induce these rules in the form of definite clauses. The final theory contained 885 clauses. For background knowledge, Progol uses a simple grammar, where the tags are terminals and predicates such as nounp (noun phrase) are non-terminals. Progol was altered to allow the caching of information about clauses generated during the induction process which greatly increased efficiency. The system achieved a per-word accuracy of 96.4% on known words drawn from sentences without quotation marks. This is on a par with other tagging systems induced from the same data [5, 2, 4] which all have accuracies in the range 96–97%. The per-sentence accuracy was 4 49.5%.