A Deep Learning Architecture for Corpus Creation for Telugu Language

A Deep Learning Architecture for Corpus Creation for Telugu Language
复制标题

用于泰卢固语语料库创建的深度学习架构

DOI:
10.1007/978-981-15-4029-5_1
复制
发表时间:
2020
期刊:
Advances in intelligent systems and computing
影响因子:
--
通讯作者:
Dhana L. Rao, Venkatesh R.
Dhana L. Rao, Venkatesh R.
中科院分区:
--
文献类型:
--
作者:
Dhana L. Rao, Venkatesh R.

文献摘要

参考文献

相似文献

由于英语作为万维网(WWW)语言的主导地位、全球化经济、社会经济和政治因素,许多自然语言正在衰落。计算语言学为保护和推广自然语言提供了前所未有的机会。然而,语料库的可用性对于利用计算语言学技术至关重要。只有少数语言拥有不同类型的语料库,而从机器可读语料库的可用性来看,大多数语言资源匮乏。泰卢固语就是这样一种语言,它是印度南部两个邦的官方语言。在本文中,我们概述了评估语言活力/危险的技术,描述了开发泰卢固语语料库的现有资源,讨论了我们开发语料库的方法,并提出了初步结果。
Many natural languages are on the decline due to the dominance of English as the language of the World Wide Web (WWW), globalized economy, socioeconomic, and political factors. Computational Linguistics offers unprecedented opportunities for preserving and promoting natural languages. However, availability of corpora is essential for leveraging the Computational Linguistics techniques. Only a handful of languages have corpora of diverse genre while most languages areresource-poorfrom the perspective of the availability of machine-readable corpora. Telugu is one such language, which is the official language of two southern states in India. In this paper, we provide an overview of techniques for assessing language vitality/endangerment, describe existing resources for developing corpora for the Telugu language, discuss our approach to developing corpora, and present preliminary results.
非洲语言濒危的意识形态和类型
DOI: --
发表时间: 2015
期刊:
影响因子: --
作者:
Friederike Lüpke
通讯作者: Friederike Lüpke
OLAC 开放语言档案社区
DOI: --
发表时间: 2004
期刊:
影响因子: --
作者:
M. Wynne
通讯作者: M. Wynne
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者:
Anita C. Faul
通讯作者: Anita C. Faul
世界语言结构地图集 (WALS)
DOI: --
发表时间: 2006
期刊:
影响因子: --
作者:
H. Hughes
通讯作者: H. Hughes
DOI: 10.2307/330061
发表时间: 1991
影响因子: 0.6
作者:
J. Fishman
通讯作者: J. Fishman