An Open Source Toolkit for Quantitative Historical Linguistics

An Open Source Toolkit for Quantitative Historical Linguistics
复制标题

定量历史语言学的开源工具包

DOI:
10.5167/uzh-84667
复制
发表时间:
2013
期刊:
影响因子:
0.7
通讯作者:
Steven Moran
Steven Moran
中科院分区:
人文科学4区
文献类型:
--
作者:
Johann;Steven Moran

文献摘要

被引文献

相似文献

鉴于历史语言学中计算和定量方法的日益增长的兴趣和发展,学者们有一个记录,测试,评估和共享复杂工作流程的基础是很重要的。我们提出了一个新的开源工具包,在历史语言学的定量任务,提供这些功能。该工具包还充当现有软件包和常用数据格式之间的接口,并在同构框架内提供新算法和现有算法的实现。我们用一个示例性的工作流程来说明该工具包的功能,该工作流程从原始语言数据开始,以自动计算的语音对齐、同源词和借词结束。然后,我们说明了与工具包提供的黄金标准数据集的评估指标。
Given the increasing interest and development of computational and quantitative methods in historical linguistics, it is important that scholars have a basis for documenting, testing, evaluating, and sharing complex workflows. We present a novel open-source toolkit for quantitative tasks in historical linguistics that offers these features. This toolkit also serves as an interface between existing software packages and frequently used data formats, and it provides implementations of new and existing algorithms within a homogeneous framework. We illustrate the toolkit’s functionality with an exemplary workflow that starts with raw language data and ends with automatically calculated phonetic alignments, cognates and borrowings. We then illustrate evaluation metrics on gold standard datasets that are provided with the toolkit.