Finding Ideographic Representations of Japanese Names Written in Latin Script via Language Identification and Corpus Validation

Finding Ideographic Representations of Japanese Names Written in Latin Script via Language Identification and Corpus Validation
复制标题

通过语言识别和语料库验证查找用拉丁文字书写的日语名字的表意表示

DOI:
10.3115/1218955.1218979
复制
发表时间:
2004
期刊:
--
影响因子:
--
通讯作者:
G. Grefenstette
G. Grefenstette
中科院分区:
--
文献类型:
--
作者:
Yan Qu;G. Grefenstette

文献摘要

被引文献

相似文献

多语言应用经常涉及处理专有名词,但在双语词典中经常缺少名称。对于涉及拉丁脚本语言和亚洲语言(例如中文、日文和韩文(CJK))之间的翻译的应用程序,该问题更加严重,其中简单的字符串复制不是解决方案。我们提出了一种新的方法来产生表意文字表示的CJK名字写在拉丁字母。所提出的方法包括首先确定名称的起源,然后使用语言特定的映射将名称回译为所有可能的中文字符。为了减少大量的计算的可能性,我们应用了一个三层过滤过程,首先通过一组经证明的二元组进行过滤,然后通过一组经证明的条款,最后通过WWW进行最终验证。我们用英日回译来说明这种方法。对日本人的名字和姓氏的测试集,我们已经取得了73%和90%,分别平均精度。
Multilingual applications frequently involve dealing with proper names, but names are often missing in bilingual lexicons. This problem is exacerbated for applications involving translation between Latin-scripted languages and Asian languages such as Chinese, Japanese and Korean (CJK) where simple string copying is not a solution. We present a novel approach for generating the ideographic representations of a CJK name written in a Latin script. The proposed approach involves first identifying the origin of the name, and then back-transliterating the name to all possible Chinese characters using language-specific mappings. To reduce the massive number of possibilities for computation, we apply a three-tier filtering process by filtering first through a set of attested bigrams, then through a set of attested terms, and lastly through the WWW for a final validation. We illustrate the approach with English-to-Japanese back-transliteration. Against test sets of Japanese given names and surnames, we have achieved average precisions of 73% and 90%, respectively.