Creation of a Doctor-Patient Dialogue Corpus Using Standardized Patients

Creation of a Doctor-Patient Dialogue Corpus Using Standardized Patients
复制标题

使用标准化患者创建医患对话语料库

DOI:
--
复制
发表时间:
2004
期刊:
--
影响因子:
--
通讯作者:
S. Ganjavi
S. Ganjavi
中科院分区:
--
文献类型:
--
作者:
Robert S. Melvin;Win May;Shrikanth S. Narayanan;P. Georgiou;S. Ganjavi

文献摘要

被引文献

相似文献

在本文中,我们描述了一个医生与病人的对话语料库,以支持语音到语音的机器翻译工作的英语-波斯语的医疗对话的发展。该语料库是通过记录和转录医学生和标准化患者(受过训练的演员,以描绘疾病或伤害受害者)之间的英语到英语对话,然后翻译成波斯语而开发的。我们将讨论以这种方式创建语料库的一些优点和缺点。好处包括能够以一种对实际医患数据不可行的方式定制语料库,并避免隐私和法律的问题,而缺点包括波斯语并非源于语音,而是英语语音的文本翻译。我们解决的问题,如对话的真实性和这些数据的系统开发的价值。
In this paper we describe the development of a doctor-patient dialogue corpus to support a speech-to-speech machine translation effort for English-Persian medical dialogues. The corpus was developed by recording and transcribing English-to-English dialogues between medical students and standardized patients (actors who have been trained to portray illness or injury victims), and then translated into Persian. We discuss some of the benefits and drawbacks to creating a corpus in this way. Benefits include the ability to customize the corpus in a way that would be infeasible for actual doctor-patient data and avoidance of privacy and legal issues, while drawbacks include the fact that the Persian does not originate as speech, but as text translation of English speech. We address concerns such as the authenticity of the dialogues and the value of such data for system development.