Borders and boundaries in Bosnian, Croatian, Montenegrin and Serbian: Twitter data to the rescue

Borders and boundaries in Bosnian, Croatian, Montenegrin and Serbian: Twitter data to the rescue
复制标题

波斯尼亚语、克罗地亚语、黑山语和塞尔维亚语的边界和边界:Twitter 数据的救援

DOI:
--
复制
发表时间:
2018
期刊:
Journal of Linguistic Geography
影响因子:
--
通讯作者:
T. Samardžić
T. Samardžić
中科院分区:
--
文献类型:
--
作者:
N. Ljubešić;M. Petrović;T. Samardžić

文献摘要

被引文献

相似文献

在本文中,我们研究了波斯尼亚语、克罗地亚语、黑山语和塞尔维亚语之间已知的 16 种语言特征的空间分布。我们对 2013 年中至 2016 年底期间收集的地理编码 Twitter 状态消息数据集进行分析。我们执行两种类型的分析。第一个通过核密度估计平滑技术找到语言变量水平空间分布的边界。然后将这些边界绘制在州边界上以进行视觉比较。第二个分析涉及国家之间的语言距离。语言变量和国家的分组是根据州边界以及每个州内 16 个变量的分布之间的 Jensen-Shannon 散度来计算的。该分析是通过衡量每个国家的变量一致性来完成的。这些分析旨在显示当前国家边界与语言边界的对应程度。他们认为克罗地亚和塞尔维亚仍然代表两个极端,反映了规范分歧的历史,而波斯尼亚和黑塞哥维那和黑山则根据变量倾向于其中一方。
In this paper we deal with the spatial distribution of 16 linguistic features known to vary between Bosnian, Croatian, Montenegrin, and Serbian. We perform our analyses on a dataset of geo-encoded Twitter status messages collected in the period from mid-2013 to the end of 2016. We perform two types of analyses. The first one finds boundaries in the spatial distribution of the linguistic variable levels through the kernel density estimation smoothing technique. These boundaries are then plotted over the state borders for a visual comparison. The second analysis deals with linguistic distance between the states. The groupings of linguistic variables and countries are calculated given the state borders and the Jensen-Shannon divergence between distributions of the 16 variables within each state. This analysis is completed with a measure of variable consistency for each country. These analyses are intended to show the extent to which current state borders correspond to linguistic boundaries. They suggest that Croatia and Serbia still represent the two extremes, reflecting a history of normative divergences, while Bosnia-Herzegovina and Montenegro, depending on the variable, lean to one or the other side.