What is a word, What is a sentence? Problems of Tokenization

What is a word, What is a sentence? Problems of Tokenization
复制标题

什么是词,什么是句子?

DOI:
--
复制
发表时间:
1994
期刊:
影响因子:
--
通讯作者:
P. Tapanainen
P. Tapanainen
中科院分区:
--
文献类型:
--
作者:
G. Grefenstette;P. Tapanainen

文献摘要

被引文献

相似文献

任何对自由出现的文本的语言学处理都必须对什么被认为是标记提供一个答案。在阿尔蒂(arti:12)社会语言中,对被认为是一个标记的东西的定义可以被精确地、明确地定义。另一方面,自然语言显示出如此丰富的多样性,以至于有许多方法可以决定什么将被视为文本计算方法的单元。在这里,我们将讨论标记化作为计算词典学的一个问题。我们的讨论将涵盖通常被认为是文本预处理的各个方面,以便为某些自动化处理做好准备。我们介绍了标记化的作用,标记化的方法,识别首字母缩略词,缩写词和正则表达式(如数字和日期)的语法。我们提出了遇到的问题,并讨论了看似无辜的选择的影响。
Any linguistic treatment of freely occurring text must provide an answer to what is considered as a token. In arti(cid:12)cial languages, the de(cid:12)nition of what is considered as a token can be precisely and unambiguously de(cid:12)ned. Natural languages, on the other hand, display such a rich variety that there are many ways to decide upon what will be considered as a unit for a computational approach to text. Here we will discuss tokenization as a problem for computational lexicography. Our discussion will cover the aspects of what is usually considered preprocessing of text in order to prepare it for some automated treatment. We present the roles of tokenization, methods of tokenizing, grammars for recognizing acronyms, abbreviations, and regular expressions such as numbers and dates. We present the problems encountered and discuss the e(cid:11)ects of seemingly innocent choices.