A European database of descriptors of English electronic texts

A European database of descriptors of English electronic texts
复制标题

欧洲英语电子文本描述符数据库

DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
J. Tyrkkö
J. Tyrkkö
中科院分区:
--
文献类型:
--
作者:
Hans;H. D. Smet;J. Tyrkkö

文献摘要

被引文献

相似文献

《信使》的读者几乎不需要被告知电子语料库在语言学(包括历史语言学)中的重要性和影响。《信使报》的第一期(实际上是“零期”)刊登了一篇关于赫尔辛基语料库汇编的文章。[4]尽管当时已经有一些语料库在使用,但赫尔辛基语料库是在英语历史研究中使用语料库的先驱。从一开始,语料库编辑的一个关键问题就是语料库的组成问题。第一个电子英语语料库,布朗语料库(1961年),使用了一套文本类别,随后被许多语料库采用和改编,这些语料库遵循相同的模式,从LOB到FLOB,皱眉等等。由马蒂里萨宁领导的赫尔辛基小组创建了一组22个“文本参数”,其不仅给出了像“文本类型”和“原型文本类别”这样的文本描述符而且还包括日期,这些因素包括文本的方言/地区出处、作者的性别和社会地位、他或她与听众或读者的关系、它与外语原文和口语的关系(Kyto 1991:x-xi)。值得注意的是,通常被认为是语料库组成的基本要求的“平衡和多样性”标准,在历史语料库中更难满足。不仅很难提供足够的原始数据,特别是早期的数据,而且文本类型和体裁也会随着时间的推移而变化和发展。这反过来又影响了语料库的使用方式,因为无论何时,只要某一特定体裁的一个世纪的作品被少数摘录所代表,就只能在非常试探性的基础上得出结论--赫尔辛基语料库的编纂者公开承认这一事实。因此,基于语料库证据的研究结果有可能以学者们甚至没有意识到的方式被倾斜。我们认为,应该进一步培养这种意识。另一个问题是数量。赫尔辛基语料库包含了大约150万个单词,旨在代表从最早的现存记录到大约
Readers of the Messenger need hardly be told of the importance and impact of electronic corpora in linguistics, including historical linguistics. The very first issue (the “Zero Issue” in fact) of the Messenger carried an article on the compilation of the Helsinki Corpus.4 Although a number of corpora were already in use at the time, the Helsinki Corpus was a pioneer when it came to the use of corpora in the historical study of English. From the beginning, a key problem for corpus compilers has been the question of corpus composition. The first electronic corpus of English, the Brown Corpus (1961), used a set of textual categories which were subsequently adopted and adapted by numerous corpora compiled following the same model from LOB to FLOB, Frown and so many others.5 Considering the rather different needs of the historical linguist, the Helsinki team led by Matti Rissanen created a set of 22 “textual parameters” which not only give textual descriptors like “Text Types” and “Prototypical Text Categories” but also includes the date, dialectal / regional provenance of the text, the author’s sex and social rank, his or her relationship with the audience or readership, its relationship to foreign-language originals and to the spoken language (Kyto 1991: x-xi). Note that the criteria of ‘balance and diversity’,6 usually considered essential requirements of corpus composition, are much harder to satisfy in a historical corpus. Not only is it often difficult to provide sufficient primary data as it is, particularly of the earlier periods, but text types and genres also change and develop over time. This in turn affects the way corpora can be used, for whenever a century of writing in any given genre is represented by a small handful of extracts, conclusions can only be drawn on a very tentative basis — a fact the compilers of the Helsinki Corpus openly acknowledge. As a consequence, findings based on corpus evidence run the risk of being slanted in ways that scholars may not even be aware of. We feel that this awareness ought to be fostered more. Another problem is quantity. The Helsinki Corpus contains some 1.5 million words intended to represent the whole history of English from the earliest extant records to about