Using Sequence Similarity Networks to Identify Partial Cognates in Multilingual Wordlists

Using Sequence Similarity Networks to Identify Partial Cognates in Multilingual Wordlists
复制标题

使用序列相似性网络识别多语言单词列表中的部分同源词

DOI:
--
复制
发表时间:
2016
期刊:
Annual Meeting of the Association for Computational Linguistics
影响因子:
--
通讯作者:
Eric Bapteste
Eric Bapteste
中科院分区:
--
文献类型:
--
作者:
Johann;P. Lopez;Eric Bapteste

文献摘要

参考文献

被引文献

相似文献

历史语言学中越来越多的数字数据要求开发自动方法来检测跨语言的同源词。最近发展的方法在中等时间深度的语言家族上工作得很好,但它们不能识别仅部分相关的单词中的同源语素。然而,部分认知是一种经常出现的现象,特别是在派生词法丰富的语言家族中。本文提出了一种部分同源词检测的试点方法,该方法用网络来表示词性之间的相似性,并借助最新的网络划分算法识别同源语素。该方法在一个新创建的基准数据集上进行了测试,数据来自汉藏语的三个分支,产生了非常有希望的结果,性能优于所有对部分认知不敏感的算法。
Increasing amounts of digital data in historical linguistics necessitate the development of automatic methods for the detection of cognate words across languages. Recently developed methods work well on language families with moderate time depths, but they are not capable of identifying cognate morphemes in words which are only partially related. Partial cog-nacy, however, is a frequently recurring phenomenon, especially in language families with productive derivational morphology. This paper presents a pilot approach for partial cognate detection in which networks are used to represent similarities be-tween word parts and cognate morphemes are identified with help of state-of-the-art algorithms for network partitioning. The approach is tested on a newly created benchmark dataset with data from three sub-branches of Sino-Tibetan and yields very promising results, outperforming all algorithms which are not sensible to partial cognacy.
超越同源性:词语之间的历史关系及其对系统发育重建的影响
DOI: 10.1093/jole/lzw006
发表时间: 2016
影响因子: 2.6
作者:
通讯作者: --