SGER: Automatic Processing of Natural Language Code Switching
SGER: Automatic Processing of Natural Language Code Switching
批准号:
0749062
负责人:
Mona Diab
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-09-01 至 2009-02-28
中文摘要
语码转换是一种自然的语言现象,说话者将两种或两种以上的语言或方言,或同一种语言的两个或两个以上的语域混合在一起。大量的社会语言学研究致力于这种广泛而普遍的现象,并且在正式语言学中也有一些先前的工作,但迄今为止,它还没有被认为是计算语言学社区感兴趣的问题。然而,在这个全球化的时代和当前信息和网络访问的爆炸,越来越多的自发生成的语言数据来自世界各地的计算研究社区提供。这样的数据丰富的代码转换在其不同的形式,所以有一个真实的需要计算语言学家来解决代码转换作为一个中心的研究问题。这种探索性的研究工作解决了如何自动处理代码切换的问题。它检查了代码切换的不同方面,允许在对现象有清晰理解的基础上创建更好的原则性算法。主要的问题围绕切换的形态和句法约束,以及如何这些约束可以模拟计算。这项研究的成果之一是对大量数据进行了注释,这些数据显示了不同语言中的代码转换,最有可能是阿拉伯语,印地语和西班牙语。本研究旨在启动一个正式的研究代码切换的计算框架,既增加了我们的理解的现象,并开发算法处理自然语言数据,体现代码切换。
英文摘要
Code switching is a natural linguistic phenomenon in which a speaker mixes two or more languages or dialects, or two or more linguistic registers from the same language. Extensive sociolinguistic studies have been dedicated to this widespread and common phenomenon and there has been some prior work in formal linguistics, but to date it has not been considered a problem of interest to the computational linguistics community. However, in this age of globalization and the current explosion in information and web access, more and more spontaneously generated linguistic data from around the world are being made available to the computational research community. Such data abounds with code switching in its different forms, so there is a real need for computational linguists to address code switching as a central research problem. This exploratory research effort addresses the issues of how to process code switching automatically. It examines the different aspects of code switching, allowing for the creation of better-principled algorithms based on a clear understanding of the phenomenon. The main questions revolve around morphological and syntactic constraints on switching and how these constraints can be modeled computationally. One of the outcomes of this research is the annotation of significant amounts of data exhibiting code switching in different languages, most likely Arabic, Hindi and Spanish. This research aims at initiating a formal study of code switching in a computational framework, which both increases our understanding of the phenomenon, and develops algorithms for processing natural language data that manifests code switching.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CI-P: Towards the Creation of a Unified Repository for MultiLingual and CrossLingual Multiword Expressions
-
批准号:1513116
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2015
-
负责人:Mona Diab
-
依托单位:
CI-ADDO-NEW: Collaborative Research: A Repository for Annotating Multilingual Code Switched Data
-
批准号:1343530
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2013
-
负责人:Mona Diab
-
依托单位:
CI-ADDO-NEW: Collaborative Research: A Repository for Annotating Multilingual Code Switched Data
-
批准号:1205556
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2012
-
负责人:Mona Diab
-
依托单位:
Collaborative Research: CI-P: Creation of an annotated repository of multilingual and multigenre code switched data for several language pairs
-
批准号:0958440
-
项目类别:Standard Grant
-
资助金额:$7.8万
-
财政年份:2010
-
负责人:Mona Diab
-
依托单位:
海外基金