Automatic Dialect Density Estimation for African American English

Automatic Dialect Density Estimation for African American English
复制标题

DOI:
10.48550/arxiv.2204.00967
复制
发表时间:
2022-04
期刊:
--
影响因子:
--
通讯作者:
A. Johnson;K. Everson;Vijay Ravi;A. Gladney;Mari Ostendorf;Abeer Alwan
A. Johnson;K. Everson;Vijay Ravi;A. Gladney;Mari Ostendorf;Abeer Alwan
中科院分区:
其他
文献类型:
--
作者:
A. Johnson;K. Everson;Vijay Ravi;A. Gladney;Mari Ostendorf;Abeer Alwan

文献摘要

被引文献

相似文献

在本文中,我们探索了非裔美国英语方言密度的自动预测,其中方言密度被定义为话语中包含非标准方言特征的单词的百分比。我们研究了几种声学和语言建模特征,包括常用的x向量表示和比较特征集,以及从音频文件的ASR转录本和韵律信息中提取的信息。为了解决有限标记数据的问题,我们使用弱监督模型将韵律和x向量特征投影到低维任务相关表示中。然后使用XGBoost模型从这些特征中预测说话者的方言密度,并在推理过程中显示哪些是最重要的。对于给定的任务,我们评估了这些特性单独和组合的效用。这项工作不依赖于手工标记的转录本,而是在CORAAL数据库的音频片段上进行的。我们在该数据库中显示了预测的AAE语音和真实的方言密度测量之间的显著相关性,并提出这项工作作为解释和减轻语音技术偏差的工具。
In this paper, we explore automatic prediction of dialect density of the African American English (AAE) dialect, where dialect density is defined as the percentage of words in an utterance that contain characteristics of the non-standard dialect. We investigate several acoustic and language modeling features, including the commonly used X-vector representation and ComParE feature set, in addition to information extracted from ASR transcripts of the audio files and prosodic information. To address issues of limited labeled data, we use a weakly supervised model to project prosodic and X-vector features into low-dimensional task-relevant representations. An XGBoost model is then used to predict the speaker's dialect density from these features and show which are most significant during inference. We evaluate the utility of these features both alone and in combination for the given task. This work, which does not rely on hand-labeled transcripts, is performed on audio segments from the CORAAL database. We show a significant correlation between our predicted and ground truth dialect density measures for AAE speech in this database and propose this work as a tool for explaining and mitigating bias in speech technology.