Using Network Approaches to Enhance the Analysis of Cross-Linguistic Polysemies

Using Network Approaches to Enhance the Analysis of Cross-Linguistic Polysemies
复制标题

使用网络方法加强跨语言多义词的分析

DOI:
--
复制
发表时间:
2013
期刊:
International Conference on Computational Semantics
影响因子:
--
通讯作者:
M. Urban
M. Urban
中科院分区:
--
文献类型:
--
作者:
Johann;A. Terhalle;M. Urban

文献摘要

被引文献

相似文献

长期以来,人们一直注意到跨语言重复出现的多义词可以作为概念关系的指标,并在最近的过去已经提出了相当多的方法来建模和分析这样的数据。虽然-考虑到数据的性质-似乎很自然地借助网络技术对其进行建模和分析,但只有少数方法明确使用它们。在本文中,我们展示了如何严格应用加权网络模型有助于获得更多的跨语言的多义词比可能使用的方法,只基于项目对项目的比较。在我们的研究中,我们使用了一个由1252个语义项组成的大型数据集,这些语义项被翻译成195种不同的语言,涵盖了44个不同的语系。通过分析从数据中重建的网络的社区结构,我们发现大多数概念(68%)可以分为104个由5个或更多节点组成的大社区。这些大社区几乎完全构成了概念领域中概念的有意义的分组。它们为深入分析历史语义学的各种主题提供了一个有效的起点,如同源检测,词源分析和语义重建。
Since long it has been noted that cross-linguistically recurring polysemies can serve as an indicator of conceptual relations, and quite a few approaches to model and analyze such data have been proposed in the recent past. Although – given the nature of the data – it seems natural to model and analyze it with the help of network techniques, there are only a few approaches which make explicit use of them. In this paper, we show how the strict application of weighted network models helps to get more out of cross-linguistic polysemies than would be possible using approaches that are only based on item-to-item comparison. For our study we use a large dataset consisting of 1252 semantic items translated into 195 different languages covering 44 different language families. By analyzing the community structure of the network reconstructed from the data, we find that a majority of the concepts (68%) can be separated into 104 large communities consisting of five and more nodes. These large communities almost exclusively constitute meaningful groupings of concepts into conceptual fields. They provide a valid starting point for deeper analyses of various topics in historical semantics, such as cognate detection, etymological analysis, and semantic reconstruction.