Spelling errors and keywords in born-digital data: a case study using the Teenage Health Freak Corpus

Spelling errors and keywords in born-digital data: a case study using the Teenage Health Freak Corpus
复制标题

DOI:
10.3366/cor.2014.0055
复制
发表时间:
2014-11-01
期刊:
影响因子:
0.5
通讯作者:
Mullany, Louise
Mullany, Louise
中科院分区:
其他
文献类型:
--
作者:
Smith, Catherine;Adolphs, Svenja;Mullany, Louise

文献摘要

被引文献

相似文献

现在以数字形式提供的语言数据丰富,以及用于数字通信的不同语言种类的增加,意味着未来不标准拼写和拼写错误的问题可能会对语料库的编纂者变得更加突出。本文通过对一个数字化语料库中关键词的拼写变异进行研究,旨在探讨这种变异的程度和影响,为今后的语料库研究提供参考。在这项研究中使用的语料库由青少年发送到健康网站的健康问题的电子邮件。关键词的生成使用原始版本的语料库和拼写错误更正的版本,和英国国家语料库(BNC)作为参考语料库。排名的关键字是非常相似的,因此,建议,根据研究目标,关键字可以可靠地生成,而无需任何拼写校正。
The abundance of language data that is now available in digital form, and the rise of distinct language varieties that are used for digital communication, means that issues of non-standard spellings and spelling errors are, in future, likely to become more prominent for compilers of corpora. This paper examines the effect of spelling variation on keywords in a born-digital corpus in order to explore the extent and impact of this variation for future corpus studies. The corpus used in this study consists of e-mails about health concerns that were sent to a health website by adolescents. Keywords are generated using the original version of the corpus and a version with spelling errors corrected, and the British National Corpus (BNC) acts as the reference corpus. The ranks of the keywords are shown to be very similar and, therefore, suggest that, depending on the research goals, keywords could be generated reliably without any need for spelling correction.