Field Testing the Tongues Speech-to-Speech Machine Translation System

Field Testing the Tongues Speech-to-Speech Machine Translation System
复制标题

现场测试 Tongues 语音到语音机器翻译系统

DOI:
--
复制
发表时间:
2002
期刊:
--
影响因子:
--
通讯作者:
E. Steinbrecher
E. Steinbrecher
中科院分区:
--
文献类型:
--
作者:
R. Frederking;A. Black;Ralf D. Brown;J. Moody;E. Steinbrecher

文献摘要

被引文献

相似文献

“舌头”便携式、快速开发的语音到语音机器翻译系统是专门为可部署原型的实际现场测试而开发的。在本文中,我们将描述该系统,使用美国陆军常规军官和幼稚的克罗地亚人进行的现场测试,以及对这些测试的评估。评估包括对问卷答案的分析,对系统记录日志的分析,以及作者的定性观察。测试的总体结果是,虽然该系统确实成功地帮助了翻译,但在准备好定期实地使用之前,还需要进一步发展。1. 该系统由美国陆军资助,用于支持美国陆军牧师的任务,这些牧师越来越多地被要求与当地居民打交道,通常没有人工翻译的好处。因此,它的目的是由一个训练有素的美国陆军牧师和一个完全天真和未经训练的非英语人士使用。舌头系统的架构和用户界面在很大程度上是基于外交家系统(Frederking et al., 2000)。使用的语音识别系统是开源的Sphinx II (Huang et al., 1992);翻译系统是一个EBMT/MEMT (ExampleBased MT/Multi-Engine MT)系统(Brown, 1996; Frederking and Nirenburg, 1994; Brown and Frederking, 1995),与Diplomat非常相似;合成系统是开源节日(Black et al., 1998)。虽然最初的系统专门用于演示英语和克罗地亚语之间的双向翻译,但设计也需要允许新语言的快速发展。为了确保快速发展,整个项目只允许用一个日历年的时间,包括合同安排、聘请语言专家等。总的开发工作同样受到限制:六个高级研究人员(本文的作者)提供了大约两(2)个全职人年的估计总数。除了高级工作人员外,还有兼职的克罗地亚线人、牧师和一些编程学生。我们应该注意到,用于训练系统的一些翻译数据是为Diplomat项目收集的(Frederking et al., 2000)。除了迅速发展之外,该系统不允许局限于一个有限的领域,而必须广泛覆盖。(这两个属性对于牧师们设想的活动都很重要。)由于我们要在短时间内以较少的预算构建一个覆盖范围广泛的系统,因此数据驱动的方法是唯一合理的选择。为了提供域内会话数据,我们在项目开始时安排了一些牧师在角色扮演对话中记录他们期望设备遇到的类型。幸运的是,牧师们熟悉角色扮演练习,并且都有相关的现场经验来重新表演。谈话双方都用英语进行。这些都是用头戴式麦克风以16KHz立体声(每个声道上有一个扬声器)进行数字录制的,因为这最接近最终系统的预期音频声道特性。我们总共录下了46段对话,时长从几分钟到20分钟不等。这提供了总共4.25小时的实际演讲。记录的谈话是逐字逐句地手工抄录的,并由母语为克罗地亚语的人翻译成克罗地亚语。使用英语录音对英语语音识别模型进行训练。抄本及其翻译被添加到EBMT系统的并行句示例库中。克罗地亚语翻译的一个子集由母语为克罗地亚语的人阅读,为克罗地亚语语音识别器创建数据,如其他地方所述(Black et al., 2002)。这种简单的方法似乎出人意料地足够。简单地把识别器、翻译器和合成器串在一起并不能构成一个非常有用的语音到语音翻译系统。一个好的界面是必要的,以使用户能够真正从中受益的方式使各个部分协同工作。根据早期外交家系统的经验,我们将语言界面设计为不对称的,克罗地亚语界面尽可能简单,而英语界面则处理任何必要的复杂性,因为牧师将接受使用该系统的培训和实践。即使是英语方面也不是特别复杂(参见图1)。我们包含了反向翻译功能,允许不了解目标语言的用户更好地评估翻译的质量。(我们不能使用从意义表示生成释义的方法,因为系统不使用任何意义表示。)我们还包含了一些用户要求的功能,例如内置的预先录制的克罗地亚语说明和解释(因为克罗地亚语使用者对设备和牧师的意图完全不了解),紧急关键短语(例如“不要动!),以及诸如能够修改该领域的翻译词典之类的增强功能,以便系统可以调整到更具体的任务。最终的系统在基于windows的东芝Libretto上运行,运行频率为200MHz,内存为192MB。在项目的时候(2000年),这是速度和规模的最佳组合。该系统配备了定制的触摸屏,因此说克罗地亚语的人根本不需要打字或使用鼠标。意识到该系统可能会在非英语参与者不熟悉计算机技术的情况下使用,我们包括了一个看起来像传统电话听筒的麦克风/扬声器听筒。它的优点是提供了一个近距离说话的麦克风,从而使语音识别更容易,同时它的外形也为大多数人所熟悉。我们已经在其他地方提供了舌头系统发展的更详细的描述(Black et al., 2002)。我们的设计为用户纠错提供了大量的机会,努力使合作的用户能够很好地沟通,以完成没有系统(或双语人类口译员)就无法完成的重要任务,尽管当前语音识别、广泛快速发展的机器翻译和语音合成容易出错。确定我们是否达到了这样的目标需要基于任务的评估;虽然组件的错误率是有用的信息,但真正的系统级问题是是否实现了通信,以及在什么级别上进行了通信。图1:舌头用户界面
The Tongues portable, rapid-development, speech-to-speech machine translation system was developed specifically to allow a realistic field-test of a deployable prototype. In this paper we will describe the system, its field-testing using regular US Army officers and naive Croatians, and the evaluation of these tests. The evaluation includes analysis of answers to a questionnaire, analysis of system transcript logs, and the authors’ qualitative observations. The overall result of the test was that while the system did successfully aid translation, it requires further development before it would be ready for regular field use. 1. The Tongues System The Tongues system was funded by the US Army to support the mission of the US Army chaplains, who are increasingly called upon to deal with local populations, usually without the benefit of human translators. It is thus intended to be used by a trained US Army chaplain with a completely naive and untrained non-English speaker. The architecture and user interface of the Tongues system were based in large measure on the Diplomat system (Frederking et al., 2000). The speech recognition system used was the open-source Sphinx II (Huang et al., 1992); the translation system was a EBMT/MEMT (ExampleBased MT/Multi-Engine MT) system (Brown, 1996; Frederking and Nirenburg, 1994; Brown and Frederking, 1995) very similar to that in Diplomat; and the synthesis system was the open-source Festival (Black et al., 1998). While the initial system was specifically to demonstrate translation in both directions between English and Croatian, the design was also required to allow rapid development for new languages. To ensure rapid development, the entire project was only allowed to take one calendar year, including contractual arrangements, hiring language experts, etc. The total development effort was similarly restricted: six senior research personnel (the authors of this paper) provided an estimated total of about two (2) fulltime person-years of effort. In addition to the senior staff, there were also part-time Croatian informants, chaplains, and some student programmers. We should note that some of the translation data used to train the system was collected for the Diplomat project (Frederking et al., 2000). In addition to rapid development, the system was not permitted to be restricted to a narrowly-limited domain, but had to be wide-coverage. (Both of these properties were important for the chaplains’ envisioned activities.) Since we were to build a broad-coverage system in a short period of time on a small budget, data-driven approaches were the only reasonable choice. In order to provide in-domain conversational data, we arranged at the start of the project to record a number of chaplains in role-playing conversations of the type they expected the device to encounter. Fortunately, the chaplains were familiar with role-playing exercises, and all had relevant field experiences to re-enact. Both sides of the conversations were spoken in English. These were digitally recorded with head-mounted microphones at 16KHz in stereo (one speaker on each channel), as this was closest to the intended audio channel characteristics of the eventual system. In all, we recorded 46 conversations, ranging from a few minutes to 20 minutes length. This provided a total of 4.25 hours of actual speech. The recorded conversations were hand-transcribed at the word level, and translated into Croatian by native Croatian speakers. The English recordings were used for training the English speech recognition models. The transcripts and their translations were added to the EBMT system’s example base of parallel sentences. A subset of the Croatian translations were read by native Croatian speakers to create data for the Croatian speech recognizer, as described elsewhere (Black et al., 2002). This simple approach appears to be surprisingly adequate. Simply stringing together a recognizer, translator, and synthesizer does not make a very useful speech-to-speech translation system. A good interface is necessary to make the parts work together in such a way that a user can actually derive benefit from it. Using our experience from the earlier Diplomat system, we designed the Tongues interface to be asymmetric, with the Croatian side being as simple as possible, and any necessary complexity handled on the English side, since the chaplain would be trained and practiced in using the system. Even the English side was not terribly complex (see Figure 1). We included a back-translation capability, to allow a user with no knowledge of the target language to better assess the quality of the translation. (We could not use the approach of generating paraphrases from meaning representations, since the system does not use any meaning representations.) We also included several user-requested features, such as built-in pre-recorded instructions and explanations for the Croatian (since the Croatian speaker is completely naive regarding the device and the chaplain’s intentions), emergency key phrases (such as “Don’t move!”), and enhancements such as being able to modify the translation lexicon in the field, so that the system could be tuned to more specific tasks. The final system ran on a Windows-based Toshiba Libretto, running at 200MHz with 192MB of memory. At the time of the project (2000) this was the best combination of speed and size that was readily available. The system was equipped with a custom touchscreen, so that the Croatian-speaker would not need to type or use a mouse at all. Aware that the system might be used in situations where the non-English participant would be unfamiliar with computer technology, we included a microphone/speaker handset that looks like a conventional telephone handset. This has the advantage of provided a close-talking microphone, thus making speech recognition easier, while coming in a form factor that will be familiar to most people. We have provided a more detailed description of the development of the Tongues system elsewhere (Black et al., 2002). Our design provides abundant opportunities for user error correction, in an effort to enable cooperative users to communicate well enough to accomplish significant tasks that they could not accomplish without the system (or a bilingual human interpreter), despite the error-prone nature of current speech recognition, broad-coverage rapiddevelopment machine translation, and speech synthesis. Determining whether we have met such a goal requires task-based evaluation; while error rates of components are useful information, the real system-level issue is whether communication is achieved, and at what level of effort. Figure 1: Tongues User Interface.