FAI: A New Paradigm for the Evaluation and Training of Inclusive Automatic Speech Recognition
FAI: A New Paradigm for the Evaluation and Training of Inclusive Automatic Speech Recognition
批准号:
2147350
负责人:
Mark Hasegawa-Johnson
金额:
$50.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-02-15 至 2025-01-31
中文摘要
自动语音识别可以通过一些小的方式提高您的工作效率:与使用图形用户界面搜索歌曲、产品或地址相比,使用自动语音识别通常可以更快地完成这些任务。 然而,对于许多人群来说,语音识别效果不太好,可能是因为地区口音,或者因为第二语言口音,或者因为残疾。 这个AI项目的公平性定义了一种思考语音技术的新方式。 在这种新的思维方式中,除非自动语音识别器对所有用户都能很好地工作,包括具有地区口音、第二语言口音和严重残疾的用户。 有三个子项目。 第一个子项目将创建语音技术研究人员可以用来测试他们的语音识别器的黑盒测试标准,以测试他们的语音识别器对不同人群的有用程度。 例如,如果研究人员发现他们的产品对某些人很有效,但对另一些人却不起作用,那么研究人员将有机会收集更多的训练数据,并进行更多的开发,以确保服务不足的社区得到更好的服务。 第二个子项目将创建玻璃盒测试标准,研究人员可以使用它来调试包容性问题。 例如,如果一个语音识别器在某种方言上有问题,那么玻璃盒方法将识别出该方言中使识别器感到困惑的特定语音,以便研究人员可以更有效地解决问题。 第三个子项目将创建训练语音识别器的新方法,以保证它对可用数据中表示的所有不同群体都同样有效。 数据将来自播客和互联网。 发言者只有在宣布自己是某一特定团体的成员时,才被确认为该团体的成员。 所有开发的软件都将以开放源代码的方式分发。自动语音识别具有使信息流民主化的潜力:人工智能对话代理可以为那些不知道去哪里寻找信息的人提供信息。在过去的五十年里,语音开发者社区对最小错误率的不懈关注已经产生了一种生产力工具,对于那些语音模式与其训练数据相匹配的人来说效果非常好:通常,受过大学教育的标准化方言的第一语言使用者,很少或没有语言障碍。 然而,对于许多人群来说,语音识别效果不太好,可能是因为他们的语音模式与标准方言有很大不同(例如,因为地区口音),因为组内异质性(例如,区域性非洲裔美国人方言),或者因为组中每个个体的语音模式表现出可变性(例如,有严重残疾的人或第二语言学习者)。 该建议的目的是创建一个新的范例,包括自动语音识别器的评估和培训。 所提出的新的评估和训练范例包括三个部分:(1)“黑盒评估”是一种评估,可以通过观察其输出来衡量语音识别器的包容性程度,而无需访问源代码或训练参数。 通过适当平衡的测试数据,统计测试可以确定系统是否为所有用户组提供相同的错误率,如果不同的组获得不同的错误率,则差异的大小可以被读取为问题大小的度量。 (2)“玻璃盒评估”是一种评估,它识别出在组之间始终区分的错误模式,并在声学信号和网络的训练参数中搜索这些错误的原因。(3)包容性优化是一系列端到端神经网络训练标准,以及训练数据集设计和增强标准,这些标准明确地平衡了对低平均错误率的需求与对低组间和说话者间方差的需求。 为了开发这些新的评估和培训模式,研究人员建议开发和分发开源数据和工具。 数据将从包括10万播客语料库在内的大型公共数据源中提取;研究人员将在语料库中搜索对话行为,其中发言者将自己与特定群体联系起来,然后将发现的群体身份和手动翻译作为开源元数据分发。 工具将使用包括K2在内的开源工具包来实现,这些工具将作为开源系统配方分发。 语音技术开发人员是一群竞争激烈的人:如果有一个数字描述了语音识别器的包容性,并且有理由相信这个数字是有科学依据的,那么世界各地的研究人员都会竞相使他们的系统更具包容性。 拟议的研究将开发这些指标和相关数据,并将其开源部署。 这项研究将在研究人员正在进行的推广计划中作为人工智能社会影响的典范。该奖项反映了NSF的法定使命,并通过使用基金会的智力价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Automatic speech recognition can improve your productivity in small ways: rather than searching for a song, a product, or an address using a graphical user interface, it is often faster to accomplish these tasks using automatic speech recognition. For many groups of people, however, speech recognition works less well, possibly because of regional accents, or because of second-language accent, or because of a disability. This Fairness in AI project defines a new way of thinking about speech technology. In this new way of thinking, an automatic speech recognizer is not considered to work well unless it works well for all users, including users with regional accents, second-language accents, and severe disabilities. There are three sub-projects. The first sub-project will create black-box testing standards that speech technology researchers can use to test their speech recognizers, in order to test how useful their speech recognizer will be for different groups of people. For example, if a researcher discovers that their product works well for some people, but not others, then the researcher will have the opportunity to gather more training data, and to perform more development, in order to make sure that the under-served community is better-served. The second sub-project will create glass-box testing standards that researchers can use to debug inclusivity problems. For example, if a speech recognizer has trouble with a particular dialect, then glass-box methods will identify particular speech sounds in that dialect that are confusing the recognizer, so that researchers can more effectively solve the problem. The third sub-project will create new methods for training a speech recognizer in order to guarantee that it works equally well for all of the different groups represented in available data. Data will come from podcasts and the Internet. Speakers will be identified as members of a particular group if and only if they declare themselves to be members of that group. All of the developed software will be distributed open-source.Automatic speech recognition has the potential to democratize the flow of information: artificially intelligent dialog agents can provide information to people who would otherwise not know where to look. The speech developer community's relentless focus on minimum error rate over the past fifty years has resulted in a productivity tool that works extremely well for those of whose speech patterns match its training data: typically, college-educated first-language speakers of a standardized dialect, with little or no speech disability. For many groups of people, however, speech recognition works less well, possibly because their speech patterns differ significantly from the standard dialect (e.g., because of regional accent), because of intra-group heterogeneity (e.g., regional African American dialects), or because the speech pattern of each individual in the group exhibits variability (e.g., people with severe disabilities, or second-language learners). The aim of this proposal is to create a new paradigm for the evaluation and training of inclusive automatic speech recognizers. The proposed new evaluation and training paradigm consists of three components: (1) A "black-box evaluation" is an evaluation that can measure the degree of inclusivity of a speech recognizer by observing its outputs, without access to source code or trained parameters. With appropriately balanced test data, a statistical test can determine whether or not a system provides all groups of users with the same error rates, and if different groups get different error rates, then the size of the difference can be read as a measurement of the size of the problem. (2) A "glass-box evaluation" is an evaluation that identifies error patterns that consistently differentiate between groups, and searches for the causes of those errors in the acoustic signal and in the trained parameters of the network. (3) Inclusive optimization is a family of end-to-end neural network training criteria, and training dataset design and augmentation criteria, that explicitly balance the need for low average error rate against the need for low inter-group and inter-speaker variance. In order to develop these new evaluation and training paradigms, the researchers propose to develop and distribute open-source data and tools. Data will be drawn from large public data sources including the 100,000-podcast corpus; researchers will search the corpus for dialog acts in which speakers identify themselves with a particular group, then distribute discovered group identities and manual transcriptions as open-source metadata. Tools will be implemented using open source toolkits including K2, and those tools will be distributed as open-source system recipes. Speech technology developers are a competitive bunch: if there is a single number that describes the inclusivity of a speech recognizer, and if there is reason to believe that number to be scientifically well-founded and desirable, then researchers all over the world will compete to make their systems more inclusive. Proposed research will develop such metrics, and associated data, and will deploy them open-source. This research will be held up as a model of the social impact of artificial intelligence in the ongoing outreach programs of the investigators.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Model-Based Fairness Metric for Speaker Verification
用于说话者验证的基于模型的公平性度量
DOI:
10.1109/asru57964.2023.10389804
发表时间:
2023
期刊:
2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU
影响因子:
--
作者:
[Jahan, Maliha, Moro-Velazquez, Laureano, Thebaud, Thomas, Dehak, Najim, Villalba, Jesús]
通讯作者:
Villalba, Jesús
RI: Small: Collaborative Research: Automatic Creation of New Speech Sound Inventories
-
批准号:1910319
-
项目类别:Standard Grant
-
资助金额:$25.98万
-
财政年份:2019
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
EAGER: Matching Non-Native Transcribers to the Distinctive Features of the Language Transcribed
-
批准号:1550145
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2015
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
FODAVA-Partner: Visualizing Audio for Anomaly Detection
-
批准号:0807329
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2008
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
RI Medium: Audio Diarization - Towards Comprehensive Description of Audio Events
-
批准号:0803219
-
项目类别:Standard Grant
-
资助金额:$24.99万
-
财政年份:2008
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
Audiovisual Distinctive-Feature-Based Recognition of Dysarthric Speech
-
批准号:0534106
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
Prosodic, Intonational, and Voice Quality Correlates of Disfluency
-
批准号:0414117
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2004
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
CAREER: Landmark-Based Speech Recognition in Music and Speech Backgrounds
-
批准号:0132900
-
项目类别:Continuing Grant
-
资助金额:$39.58万
-
财政年份:2002
-
负责人:Mark Hasegawa-Johnson
-
依托单位:
海外基金