An algorithm for suffix stripping

An algorithm for suffix stripping
复制标题

DOI:
10.1108/00330330610681286
复制
发表时间:
2006-01-01
影响因子:
--
通讯作者:
Porter, M. F.
Porter, M. F.
中科院分区:
社会科学3区
文献类型:
--
作者:
Porter, M. F.

文献摘要

被引文献

相似文献

目的--自动去除英语单词中的后缀在信息检索领域具有特殊的意义。这项工作最初是在1980年出版的程序和再版的一系列文章的一部分,纪念40周年的journal.Design/方法/途径-后缀剥离的算法进行了描述,这已被实施为一个简短的,快速的程序BCPL.Findings -虽然简单,它的表现略好于一个更复杂的系统,它已被比较。它通过将复杂后缀视为由简单后缀组成的复合词来有效地工作,并在许多步骤中删除简单后缀。在每个步骤中,后缀的去除取决于剩余的词干的形式,这通常涉及其音节长度的测量。独创性/价值-这件作品提供了一个有用的历史文献信息检索。
Purpose - The automatic removal of suffixes from words in English is of particular interest in the field of information retrieval. This work was originally published in Program in 1980 and is republished as part of a series of articles commemorating the 40th anniversary of the journal.Design/methodology/approach - An algorithm for suffix stripping is described, which has been implemented as a short, fast program in BCPL.Findings - Although simple, it performs slightly better than a much more elaborate system with which it has been compared. It effectively works by treating complex suffixes as compounds made up of simple suffixes, and removing the simple suffixes in a number of steps. In each step the removal of the suffix is made to depend upon the form of the remaining stem, which usually involves a measure of its syllable length.Originality/value - The piece provides a useful historical document on information retrieval.