A European database of descriptors of English electronic texts
A European database of descriptors of English electronic texts
复制标题
欧洲英语电子文本描述符数据库
DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
J. Tyrkkö
中科院分区:
文献类型:
--
作者:
Hans;H. D. Smet;J. Tyrkkö
Readers of the Messenger need hardly be told of the importance and impact of electronic corpora in linguistics, including historical linguistics. The very first issue (the “Zero Issue” in fact) of the Messenger carried an article on the compilation of the Helsinki Corpus.4 Although a number of corpora were already in use at the time, the Helsinki Corpus was a pioneer when it came to the use of corpora in the historical study of English. From the beginning, a key problem for corpus compilers has been the question of corpus composition. The first electronic corpus of English, the Brown Corpus (1961), used a set of textual categories which were subsequently adopted and adapted by numerous corpora compiled following the same model from LOB to FLOB, Frown and so many others.5 Considering the rather different needs of the historical linguist, the Helsinki team led by Matti Rissanen created a set of 22 “textual parameters” which not only give textual descriptors like “Text Types” and “Prototypical Text Categories” but also includes the date, dialectal / regional provenance of the text, the author’s sex and social rank, his or her relationship with the audience or readership, its relationship to foreign-language originals and to the spoken language (Kyto 1991: x-xi). Note that the criteria of ‘balance and diversity’,6 usually considered essential requirements of corpus composition, are much harder to satisfy in a historical corpus. Not only is it often difficult to provide sufficient primary data as it is, particularly of the earlier periods, but text types and genres also change and develop over time. This in turn affects the way corpora can be used, for whenever a century of writing in any given genre is represented by a small handful of extracts, conclusions can only be drawn on a very tentative basis — a fact the compilers of the Helsinki Corpus openly acknowledge. As a consequence, findings based on corpus evidence run the risk of being slanted in ways that scholars may not even be aware of. We feel that this awareness ought to be fostered more. Another problem is quantity. The Helsinki Corpus contains some 1.5 million words intended to represent the whole history of English from the earliest extant records to about