SGER: Automatic Processing of Natural Language Code Switching
SGER: Automatic Processing of Natural Language Code Switching
批准号:
0749062
负责人:
Mona Diab
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-09-01 至 2009-02-28
中文摘要
语码转换是说话者将两种或两种以上的语言或方言,或同一种语言的两个或两个以上的语域混合在一起的一种自然语言现象。广泛的社会语言学研究已经致力于这一广泛而普遍的现象,并且在形式语言学中已经有了一些先前的工作,但迄今为止,它还没有被认为是计算语言学社区感兴趣的问题。然而,在这个全球化的时代,以及当前信息和网络访问的爆炸式增长,越来越多来自世界各地的自发生成的语言数据被提供给计算研究界。这些数据中充满了不同形式的代码转换,因此计算语言学家确实需要将代码转换作为一个中心研究问题来解决。这项探索性的研究工作解决了如何自动处理代码转换的问题。它检查了代码转换的不同方面,允许基于对该现象的清晰理解创建更好的原则算法。主要问题围绕着切换的形态和句法约束以及如何对这些约束进行计算建模。这项研究的成果之一是对大量数据的注释,这些数据显示了不同语言的代码转换,最有可能是阿拉伯语、印地语和西班牙语。本研究旨在在计算框架中启动对代码转换的正式研究,这既增加了我们对这一现象的理解,又开发了处理体现代码转换的自然语言数据的算法。
英文摘要
Code switching is a natural linguistic phenomenon in which a speaker mixes two or more languages or dialects, or two or more linguistic registers from the same language. Extensive sociolinguistic studies have been dedicated to this widespread and common phenomenon and there has been some prior work in formal linguistics, but to date it has not been considered a problem of interest to the computational linguistics community. However, in this age of globalization and the current explosion in information and web access, more and more spontaneously generated linguistic data from around the world are being made available to the computational research community. Such data abounds with code switching in its different forms, so there is a real need for computational linguists to address code switching as a central research problem. This exploratory research effort addresses the issues of how to process code switching automatically. It examines the different aspects of code switching, allowing for the creation of better-principled algorithms based on a clear understanding of the phenomenon. The main questions revolve around morphological and syntactic constraints on switching and how these constraints can be modeled computationally. One of the outcomes of this research is the annotation of significant amounts of data exhibiting code switching in different languages, most likely Arabic, Hindi and Spanish. This research aims at initiating a formal study of code switching in a computational framework, which both increases our understanding of the phenomenon, and develops algorithms for processing natural language data that manifests code switching.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CI-P: Towards the Creation of a Unified Repository for MultiLingual and CrossLingual Multiword Expressions
-
批准号:1513116
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2015
-
负责人:Mona Diab
-
依托单位:
CI-ADDO-NEW: Collaborative Research: A Repository for Annotating Multilingual Code Switched Data
-
批准号:1343530
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2013
-
负责人:Mona Diab
-
依托单位:
CI-ADDO-NEW: Collaborative Research: A Repository for Annotating Multilingual Code Switched Data
-
批准号:1205556
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2012
-
负责人:Mona Diab
-
依托单位:
Collaborative Research: CI-P: Creation of an annotated repository of multilingual and multigenre code switched data for several language pairs
-
批准号:0958440
-
项目类别:Standard Grant
-
资助金额:$7.8万
-
财政年份:2010
-
负责人:Mona Diab
-
依托单位:
海外基金