A Study of Spoken Audio Processing using Machine Learning for Libraries, Archives and Museums (LAM)

A Study of Spoken Audio Processing using Machine Learning for Libraries, Archives and Museums (LAM)
复制标题

DOI:
10.1109/bigdata50022.2020.9378438
复制
发表时间:
2020-12
期刊:
2020 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Weijia Xu;M. Esteva;Peter Cui;E. Castillo;Kewen Wang;H. R. Hopkins;Tanya E. Clement;Aaron Choate;Ruizhu Huang
Weijia Xu;M. Esteva;Peter Cui;E. Castillo;Kewen Wang;H. R. Hopkins;Tanya E. Clement;Aaron Choate;Ruizhu Huang
中科院分区:
其他
文献类型:
--
作者:
Weijia Xu;M. Esteva;Peter Cui;E. Castillo;Kewen Wang;H. R. Hopkins;Tanya E. Clement;Aaron Choate;Ruizhu Huang

文献摘要

被引文献

相似文献

随着在图书馆、档案馆和博物馆(LAM)中提供对口语音频收藏的访问的需求增加,对高效且一致地处理它们的需求也增加。传统上,音频处理涉及收听音频文件、进行手动转录以及应用受控主题术语来描述它们。此工作流程每次录制都需要花费大量时间。在这项研究中,我们研究了机器学习(ML)是否以及如何以符合LAM最佳实践的方式促进音频集合的处理。我们使用StoryCorps收集的口述历史“Las Historias”,以及手动分配的固定主题(元数据)来描述其中的每一个。我们的方法有两个主要阶段。首先,使用两种自动语音识别(ASR)方法自动转录音频文件。接下来,我们使用转录数据和现有元数据构建不同的监督ML模型进行标签预测。在这些阶段中,对结果进行定量和定性分析。该工作流程在灵活的Web框架IDOLS中实现,以降低LAM专业人员的技术障碍。通过允许用户向超级计算机提交ML作业,复制工作流,更改配置以及透明地查看和提供反馈,该工作流允许用户与LAM专业价值观同步。该研究有几个结果,包括不同转录方法之间的质量比较以及该质量对标签预测准确性的影响。该研究还揭示了使用手动分配的元数据来构建模型的局限性,我们提出了构建成功训练数据的替代策略。
As the need to provide access to spoken word audio collections in libraries, archives, and museums (LAM) increases, so does the need to process them efficiently and consistently. Traditionally, audio processing involves listening to the audio files, conducting manual transcription, and applying controlled subject terms to describe them. This workflow takes significant time with each recording. In this study, we investigate if and how machine learning (ML) can facilitate processing of audio collections in a manner that corresponds with LAM best practices. We use the StoryCorps collection of oral histories "Las Historias," and fixed subjects (metadata) that are manually assigned to describe each of them. Our methodology has two main phases. First, audio files are automatically transcribed using two automatic speech recognition (ASR) methods. Next, we build different supervised ML models for label prediction using the transcription data and the existing metadata. Throughout these phases the results are analyzed quantitatively and qualitatively. The workflow is implemented within the flexible web framework IDOLS to lower technical barriers for LAM professionals. By allowing users to submit ML jobs to supercomputers, reproduce workflows, change configurations, and view and provide feedback transparently, this workflow allows users to be in sync with LAM professional values. The study has several outcomes including a comparison of the quality between different transcription methods and the impact of that quality on label prediction accuracy. The study also unveiled the limitations of using manually assigned metadata to build models, to which we suggest alternate strategies for building successful training data.