Extracting Biomedical Entities from Noisy Audio Transcripts

Extracting Biomedical Entities from Noisy Audio Transcripts
复制标题

DOI:
10.48550/arxiv.2403.17363
复制
发表时间:
2024-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Nima Ebadi;Kellen Morgan;Adrian Tan;Billy Linares;Sheri Osborn;Emma Majors;Jeremy Davis;Anthony Rios
Nima Ebadi;Kellen Morgan;Adrian Tan;Billy Linares;Sheri Osborn;Emma Majors;Jeremy Davis;Anthony Rios
中科院分区:
其他
文献类型:
--
作者:
Nima Ebadi;Kellen Morgan;Adrian Tan;Billy Linares;Sheri Osborn;Emma Majors;Jeremy Davis;Anthony Rios

文献摘要

相似文献

自动语音识别(ASR)技术是将口语转录为文本的基础,在临床领域有相当多的应用,包括简化医疗转录和与电子健康记录(EHR)系统集成。然而,挑战仍然存在,特别是当transmitting包含噪声时,导致应用自然语言处理(NLP)模型时性能显着下降。命名实体识别(NER)是一项重要的临床任务,特别受到这种噪声的影响,通常被称为ASR-NLP差距。先前的工作主要研究了ASR在干净录音中的效率,留下了关于噪声环境中性能的研究空白。本文介绍了一种新的数据集BioASR-NER,旨在弥合生物医学领域的ASR-NLP差距,专注于从成人电话认知简短测试(BTACT)考试中提取药物不良反应和提及的实体。我们的数据集提供了近2,000个干净和嘈杂录音的全面集合。在解决噪声的挑战,我们提出了一种创新的成绩单清洗方法使用GPT-4,调查零拍摄和少数拍摄的方法。我们的研究进一步深入研究了错误分析,揭示了转录软件中的错误类型,GPT-4的更正以及GPT-4面临的挑战。本文旨在促进对ASR-NLP差距的更好理解和潜在解决方案,最终支持增强的医疗记录实践。
Automatic Speech Recognition (ASR) technology is fundamental in transcribing spoken language into text, with considerable applications in the clinical realm, including streamlining medical transcription and integrating with Electronic Health Record (EHR) systems. Nevertheless, challenges persist, especially when transcriptions contain noise, leading to significant drops in performance when Natural Language Processing (NLP) models are applied. Named Entity Recognition (NER), an essential clinical task, is particularly affected by such noise, often termed the ASR-NLP gap. Prior works have primarily studied ASR’s efficiency in clean recordings, leaving a research gap concerning the performance in noisy environments. This paper introduces a novel dataset, BioASR-NER, designed to bridge the ASR-NLP gap in the biomedical domain, focusing on extracting adverse drug reactions and mentions of entities from the Brief Test of Adult Cognition by Telephone (BTACT) exam. Our dataset offers a comprehensive collection of almost 2,000 clean and noisy recordings. In addressing the noise challenge, we present an innovative transcript-cleaning method using GPT-4, investigating both zero-shot and few-shot methodologies. Our study further delves into an error analysis, shedding light on the types of errors in transcription software, corrections by GPT-4, and the challenges GPT-4 faces. This paper aims to foster improved understanding and potential solutions for the ASR-NLP gap, ultimately supporting enhanced healthcare documentation practices.