SGER: Automatic Processing of Natural Language Code Switching
SGER: Automatic Processing of Natural Language Code Switching
批准号:
0749062
负责人:
Mona Diab
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-09-01 至 2009-02-28
中文摘要
语码转换是说话人混合两种或两种以上语言或方言,或同一语言中两种或两种以上语域的自然语言现象。大量的社会语言学研究致力于这一普遍而普遍的现象,形式语言学也有一些前人的工作,但到目前为止,它还没有被计算语言学界认为是一个感兴趣的问题。然而,在这个全球化的时代,以及目前信息和网络访问的爆炸性增长,越来越多来自世界各地的自发生成的语言数据正在提供给计算研究界。这类数据中有大量不同形式的语码转换,因此计算语言学家确实需要将语码转换作为一个中心研究问题来解决。这项探索性的研究工作解决了如何自动处理代码转换的问题。它研究了代码切换的不同方面,允许在清楚理解该现象的基础上创建更有原则的算法。主要问题围绕着对转换的形态和句法约束,以及这些约束如何在计算上建模。这项研究的结果之一是对大量数据进行了注释,这些数据显示出不同语言的代码转换,很可能是阿拉伯语、印地语和西班牙语。这项研究的目的是在计算框架中启动对语码转换的正式研究,这既增加了我们对这一现象的理解,也为处理体现语码转换的自然语言数据开发了算法。
英文摘要
Code switching is a natural linguistic phenomenon in which a speaker mixes two or more languages or dialects, or two or more linguistic registers from the same language. Extensive sociolinguistic studies have been dedicated to this widespread and common phenomenon and there has been some prior work in formal linguistics, but to date it has not been considered a problem of interest to the computational linguistics community. However, in this age of globalization and the current explosion in information and web access, more and more spontaneously generated linguistic data from around the world are being made available to the computational research community. Such data abounds with code switching in its different forms, so there is a real need for computational linguists to address code switching as a central research problem. This exploratory research effort addresses the issues of how to process code switching automatically. It examines the different aspects of code switching, allowing for the creation of better-principled algorithms based on a clear understanding of the phenomenon. The main questions revolve around morphological and syntactic constraints on switching and how these constraints can be modeled computationally. One of the outcomes of this research is the annotation of significant amounts of data exhibiting code switching in different languages, most likely Arabic, Hindi and Spanish. This research aims at initiating a formal study of code switching in a computational framework, which both increases our understanding of the phenomenon, and develops algorithms for processing natural language data that manifests code switching.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CI-P: Towards the Creation of a Unified Repository for MultiLingual and CrossLingual Multiword Expressions
-
批准号:1513116
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2015
-
负责人:Mona Diab
-
依托单位:
CI-ADDO-NEW: Collaborative Research: A Repository for Annotating Multilingual Code Switched Data
-
批准号:1343530
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2013
-
负责人:Mona Diab
-
依托单位:
CI-ADDO-NEW: Collaborative Research: A Repository for Annotating Multilingual Code Switched Data
-
批准号:1205556
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2012
-
负责人:Mona Diab
-
依托单位:
Collaborative Research: CI-P: Creation of an annotated repository of multilingual and multigenre code switched data for several language pairs
-
批准号:0958440
-
项目类别:Standard Grant
-
资助金额:$7.8万
-
财政年份:2010
-
负责人:Mona Diab
-
依托单位:
海外基金