A First South African Corpus of Multilingual Code-switched Soap Opera Speech
A First South African Corpus of Multilingual Code-switched Soap Opera Speech
复制标题
南非第一个多语言代码转换肥皂剧语音语料库
DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
T. Niesler
中科院分区:
文献类型:
--
作者:
E. V. D. Westhuizen;T. Niesler
The corpus comprises 26.9 hours of annotated multilingual speech that contains examples of code-switching in isiZulu, isiXhosa, Setswana, Sesotho and English. The speech was obtained from South African soap operas. Code-switching between English and one of the Bantu languages is by far most prevalent in the data. Although not very common, switches between the Bantu languages themselves also occur. An initial attempt to align the audio extracted from soap opera episodes with the corresponding scripts revealed that actors very often perform ad lib. The speech and the examples of code-switching it contains can therefore be considered to be spontaneous.