PanPhon: A Resource for Mapping IPA Segments to Articulatory Feature Vectors

PanPhon: A Resource for Mapping IPA Segments to Articulatory Feature Vectors
复制标题

DOI:
--
复制
发表时间:
2016-12
期刊:
--
影响因子:
--
通讯作者:
David R. Mortensen;Patrick Littell;Akash Bharadwaj;Kartik Goyal;Chris Dyer;Lori S. Levin
David R. Mortensen;Patrick Littell;Akash Bharadwaj;Kartik Goyal;Chris Dyer;Lori S. Levin
中科院分区:
其他
文献类型:
--
作者:
David R. Mortensen;Patrick Littell;Akash Bharadwaj;Kartik Goyal;Chris Dyer;Lori S. Levin

文献摘要

被引文献

相似文献

这篇论文提供了越来越多的证据,证明当与适当的机器学习技术相结合时,语言驱动的、信息丰富的表示可以胜过语言数据的单热编码。特别是,我们表明语音特征优于基于字符的模型。 PanPhon 是一个将 5,000 多个 IPA 片段与 21 个子片段发音特征相关联的数据库。我们证明该数据库可以提高各种 NER 相关任务的性能。在语音感知方面,基于 PanPhon 特征构建的神经 CRF 模型能够比基于字符的模型在单语西班牙语和土耳其语 NER 任务上表现更好。它们也被证明在转移模型(如乌兹别克斯坦和土耳其之间)中运作良好。 PanPhon 功能还对正字法到 IPA 转换任务做出了巨大贡献。
This paper contributes to a growing body of evidence that—when coupled with appropriate machine-learning techniques–linguistically motivated, information-rich representations can outperform one-hot encodings of linguistic data. In particular, we show that phonological features outperform character-based models. PanPhon is a database relating over 5,000 IPA segments to 21 subsegmental articulatory features. We show that this database boosts performance in various NER-related tasks. Phonologically aware, neural CRF models built on PanPhon features are able to perform better on monolingual Spanish and Turkish NER tasks that character-based models. They have also been shown to work well in transfer models (as between Uzbek and Turkish). PanPhon features also contribute measurably to Orthography-to-IPA conversion tasks.