课题基金 / 基金详情

CI-ADDO-NEW: Collaborative Research: A Repository for Annotating Multilingual Code Switched Data

CI-ADDO-NEW: Collaborative Research: A Repository for Annotating Multilingual Code Switched Data
CI-ADDO-NEW:协作研究:用于注释多语言代码交换数据的存储库
批准号:
1462142
负责人:
Thamar Solorio
金额:
$24.03万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-08-31 至 2018-08-31

项目摘要

项目成果

Thamar Solorio的其他基金

相似基金

相关文献

中文摘要
翻译
语言代码转换(LCS)是在双语或多语使用者的共享语言之间来回转换的实践。这种现象在有语言界限的地理区域或有大量移民群体的地区尤为普遍。不同的语言对和/或体裁中,语言的各个层面(语音、形态、句法、语义和话语语用)都可能涉及到语言交际。当输入包括LCS时,为单一语言训练的计算算法很快就会崩溃。在计算语言学(CL)研究LCS的一个主要障碍是缺乏大型的,准确注释的LCS数据语料库。在这个项目中,收集了一个大型的LCS数据库,并开发了一个大型的注释基础设施。它始终以不同的形式(语音和文本),在不同的语言粒度级别,并在不同的语言对反映不同的语言类型(标准阿拉伯语和方言阿拉伯语,阿拉伯语-英语,西班牙语-英语,中文-英语,印地语-英语)进行注释。这一基础设施和统一的大型LCS数据资源是CL研究社区热切期待的,因为带注释的LCS数据为自适应学习算法和处理不同的数据源提供了一个自然的测试平台,以及真正的多语言处理框架。它也将有利于社会语言学和理论语言学研究者,并提供一个跨学科合作研究的平台。最后,LCS的研究有助于克服偏见,对多语言的发言者,展示了创造性的发言者在利用他们的口头剧目。这一结果对于移民人口多样化的美国的K-12教育和考试政策尤为重要。
英文摘要
Linguistic code switching (LCS) is the practice of switching back and forth between the shared languages of bilingual or multilingual speakers. This phenomenon is particularly prevalent in geographic regions with linguistic boundaries or where there are large immigrant groups. Various levels of language (phonological, morphological, syntactic, semantic and discourse-pragmatic) may be implicated in LCS in different language pairs and/or genres. Computational algorithms trained for a single language quickly break down when the input includes LCS. A major barrier to research on LCS in computational linguistics (CL) has been the lack of large, accurately annotated corpora of LCS data. In this project, a large repository of LCS data is collected and a large annotation infrastructure is developed. It is consistently annotated in different modalities (speech and text), at various levels of linguistic granularity, and across different language pairs reflecting different linguistic typologies (Standard Arabic and Dialectal Arabic, Arabic-English, Spanish-English, Chinese-English, Hindi-English). The focus of the effort is on intra-sentential LCS.This infrastructure and unified large LCS data resource is eagerly awaited by the CL research community, since annotated LCS data provides a natural test-bed for adaptive learning algorithms and the handling of diverse data sources, as well as a framework for genuine multilingual processing. It will also be of benefit to sociolinguistic and theoretical linguistic researchers, and provide a platform for collaborative interdisciplinary research. Finally, research on LCS helps overcome biases against multilingual speakers by demonstrating the creativity of such speakers in exploiting their verbal repertoires. Such a result is particularly important for K-12 education and testing policies in the USA with its diverse immigrant population.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
IRES Track I: US-Mexico Collaboration on Multimodal Detection of Objectionable Content in Online Videos in Spanish and English
  • 批准号:
    2106892
  • 项目类别:
    Standard Grant
  • 资助金额:
    $29.97万
  • 财政年份:
    2021
  • 负责人:
    Thamar Solorio
  • 依托单位:
Workshop on desiderata for a multimodal dataset for objectionable content detection
  • 批准号:
    2036368
  • 项目类别:
    Standard Grant
  • 资助金额:
    $4.4万
  • 财政年份:
    2020
  • 负责人:
    Thamar Solorio
  • 依托单位:
RI: Small: Robust Models for Sequence Labelling in Social Media Data
  • 批准号:
    1910192
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.79万
  • 财政年份:
    2019
  • 负责人:
    Thamar Solorio
  • 依托单位:
CAREER: Authorship Analysis in Cross-Domain Settings
  • 批准号:
    1462141
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $46.96万
  • 财政年份:
    2014
  • 负责人:
    Thamar Solorio
  • 依托单位:
海外基金