Rethinking End-to-End Evaluation of Decomposable Tasks: A Case Study on Spoken Language Understanding

Rethinking End-to-End Evaluation of Decomposable Tasks: A Case Study on Spoken Language Understanding
复制标题

DOI:
10.21437/interspeech.2021-1537
复制
发表时间:
2021-06
期刊:
2022 RIVF International Conference on Computing and Communication Technologies (RIVF)
影响因子:
--
通讯作者:
Siddhant Arora;Alissa Ostapenko;Vijay Viswanathan;Siddharth Dalmia;Florian Metze;Shinji Watanabe;A. Black
Siddhant Arora;Alissa Ostapenko;Vijay Viswanathan;Siddharth Dalmia;Florian Metze;Shinji Watanabe;A. Black
中科院分区:
其他
文献类型:
--
作者:
Siddhant Arora;Alissa Ostapenko;Vijay Viswanathan;Siddharth Dalmia;Florian Metze;Shinji Watanabe;A. Black

文献摘要

被引文献

相似文献

可分解的任务是复杂的,并且由子任务的层次结构组成。例如,口语意图预测结合了自动语音识别和自然语言理解。然而,现有的基准通常只为表面级别的子任务提供示例。因此,在这些基准上具有类似性能的模型可能在其他子任务上具有未观察到的性能差异。为了让有竞争力的端到端的架构之间有见地的比较,我们提出了一个框架来构建强大的测试集,使用坐标上升子任务特定的效用函数。给定一个可分解任务的数据集,我们的方法最佳地为每个子任务创建一个测试集,以单独评估端到端模型的子组件。使用口语理解作为案例研究,我们为Fluent Speech Commands和Snips SmartLights数据集生成新的拆分。每个分割有两个测试集:一个是评估自然语言理解能力的保留话语,另一个是测试语音处理技能的保留说话者。我们的分割发现端到端系统之间的性能差距高达10%,这些系统在原始测试集上彼此相差不到1%。这些性能差距允许在不同架构之间进行更现实和可操作的比较,从而推动未来的模型开发。我们为社区发布我们的拆分和工具。
Decomposable tasks are complex and comprise of a hierarchy of sub-tasks. Spoken intent prediction, for example, combines automatic speech recognition and natural language understanding. Existing benchmarks, however, typically hold out examples for only the surface-level sub-task. As a result, models with similar performance on these benchmarks may have unobserved performance differences on the other sub-tasks. To allow insightful comparisons between competitive end-to-end architectures, we propose a framework to construct robust test sets using coordinate ascent over sub-task specific utility functions. Given a dataset for a decomposable task, our method optimally creates a test set for each sub-task to individually assess sub-components of the end-to-end model. Using spoken language understanding as a case study, we generate new splits for the Fluent Speech Commands and Snips SmartLights datasets. Each split has two test sets: one with held-out utterances assessing natural language understanding abilities, and one with held-out speakers to test speech processing skills. Our splits identify performance gaps up to 10% between end-to-end systems that were within 1% of each other on the original test sets. These performance gaps allow more realistic and actionable comparisons between different architectures, driving future model development. We release our splits and tools for the community.