Functions of Code-Switching in Tweets: An Annotation Framework and Some Initial Experiments

Functions of Code-Switching in Tweets: An Annotation Framework and Some Initial Experiments
复制标题

推文中的代码切换功能:注释框架和一些初步实验

DOI:
--
复制
发表时间:
2016
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
Niloy Ganguly
Niloy Ganguly
中科院分区:
--
文献类型:
--
作者:
Rafiya Begum;Kalika Bali;M. Choudhury;Koustav Rudra;Niloy Ganguly

文献摘要

被引文献

相似文献

两种语言之间的语码转换(CS)在社会多语言的社区中非常常见,其中说话者在相互交流时在两种或多种语言之间切换。几十年来,语言学家一直在口语中广泛研究CS,但随着社交媒体和不太正式的计算机介导通信的普及,我们现在看到CS在文本形式中的使用大幅增加。这提出了有趣的挑战,需要计算处理这种代码转换的数据。与任何计算语言学分析和自然语言处理工具和应用程序一样,我们需要注释数据来理解,处理和生成代码转换语言。在本研究中,我们重点关注从印地语-英语双语者的Twitter流中提取的英语和印地语推文之间的CS。本文在语言学分析和初步实验的基础上,提出了一种用于注释印地语-英语(Hi-En)语码转换推文中CS语用功能的注释方案。
Code-Switching (CS) between two languages is extremely common in communities with societal multilingualism where speakers switch between two or more languages when interacting with each other. CS has been extensively studied in spoken language by linguists for several decades but with the popularity of social-media and less formal Computer Mediated Communication, we now see a big rise in the use of CS in the text form. This poses interesting challenges and a need for computational processing of such code-switched data. As with any Computational Linguistic analysis and Natural Language Processing tools and applications, we need annotated data for understanding, processing, and generation of code-switched language. In this study, we focus on CS between English and Hindi Tweets extracted from the Twitter stream of Hindi-English bilinguals. We present an annotation scheme for annotating the pragmatic functions of CS in Hindi-English (Hi-En) code-switched tweets based on a linguistic analysis and some initial experiments.