Language Preservation 2.0: Crowdsourcing Oral Language Documentation using Mobile Devices
Language Preservation 2.0: Crowdsourcing Oral Language Documentation using Mobile Devices
批准号:
1160639
负责人:
Mark Liberman
金额:
$10.15万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-07-01 至 2014-12-31
中文摘要
语言保存2.0这个试验项目的目的是证明一种新方法记录濒危语言的可行性。为了在一种语言不再被使用之后对它进行广泛的调查,我们需要相当于现存的圣经希伯来语文本的百万单词,或者现存的古典拉丁语的五百万单词。 但对于没有重要文化的濒危语言来说,这种规模的多样化文本收藏似乎遥不可及。假设每小时大约有10,000个单词的典型说话速度,一百小时的录音演讲-对话,叙述或口述历史-将给我们相当于一百万个单词的文本。 在社区的参与下,数百小时的录音是很容易获得的,但是,鉴于讲母语的识字人数很少,而且这种转录工作很耗时,每小时的音频可能需要200个小时的工作,转录这么大的音频集是一项艰巨的任务。我们建议通过替代复述和口头翻译来解决这个问题:一个或多个母语者重复录音中的每个短语,缓慢而仔细地说,然后将其翻译成一种记录更好的语言。翻译段落作为分析其他未知语言的一种方法的实用性已经被证明了很多次,从罗塞塔石碑开始。 这方面的任务比较容易,因为至少一般来说,我们可以得到一个语法草图。 我们在这个项目中的目标是展示复述的效用。我们相信,语言学家,从相对较少的语言知识开始,可以产生足够好的语音音译,以支持随后的分析,从而产生连贯的文本,这一过程类似于(但更容易)让前几代学者学习阅读古埃及语或苏美尔语的过程。
英文摘要
Language Preservation 2.0The purpose of this pilot project is to demonstrate the feasibility of a new approach to documenting endangered languages.To allow wide-ranging investigation of a language even after it is no longer spoken, we need the equivalent of the million words of extant biblical Hebrew texts, or the five million words of extant classical Latin. But for endangered languages without a significant culture of literacy, diverse text collections on this scale seem out of reach. Given typical speaking rates of about 10,000 word-equivalents per hour, a hundred hours of recorded speech -- conversations, narratives, or oral histories -- would give us the equivalent of a million words of text. With community involvement, hundreds of hours of such recordings are easily within reach.However, transcribing such large audio collections is a daunting task, given the small number of literate native speakers and the time-consuming nature of such transcription, which can take 200 hours of work for every hour of audio. We propose to solve this problem by substituting re-speaking and verbal translation: one or more native speakers repeats each phrase of a recording, speaking slowly and carefully, and then translates it into a better-documented language.The utility of translated passages as a way to analyze otherwise-unknown languages has been demonstrated many times, starting with the Rosetta Stone. This aspect of our task is easier, since at least a grammatical sketch will in general be available. Our goal in this project is to demonstrate the utility of re-speaking. We believe that linguists, starting out with relatively little knowledge of a language, can produce phonetic transcriptions that will be good enough to support subsequent analysis resulting in coherent texts, in a process analogous to (but easier than) the process that allowed previous generations of scholars to learn to read ancient Egyptian or Sumerian.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CI-NEW: NIEUW: Novel Incentives and Workflows in Linguistic Data Collection and Annotation
-
批准号:1730377
-
项目类别:Standard Grant
-
资助金额:$121.85万
-
财政年份:2017
-
负责人:Mark Liberman
-
依托单位:
EAGER: Mining a Year of Speech
-
批准号:1048900
-
项目类别:Standard Grant
-
资助金额:$9.99万
-
财政年份:2010
-
负责人:Mark Liberman
-
依托单位:
Prosodic Systems in New Guinea: Integrating computational and typological approaches to linguistic analysis
-
批准号:0951651
-
项目类别:Standard Grant
-
资助金额:$29.93万
-
财政年份:2010
-
负责人:Mark Liberman
-
依托单位:
Collaborative Research: OLAC: Accessing the World's Language Resources
-
批准号:0723357
-
项目类别:Continuing Grant
-
资助金额:$14.7万
-
财政年份:2007
-
负责人:Mark Liberman
-
依托单位:
ITR-SCOTUS: A Resource for Collaborative Research in Speech Technology, Linguistics, Decision Processes and the Law
-
批准号:0325739
-
项目类别:Continuing Grant
-
资助金额:$72.5万
-
财政年份:2003
-
负责人:Mark Liberman
-
依托单位:
Querying Linguistic Databases
-
批准号:0317826
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:Mark Liberman
-
依托单位:
Eletronic Materials For Natural Language Research
-
批准号:9113530
-
项目类别:Standard Grant
-
资助金额:$13.99万
-
财政年份:1991
-
负责人:Mark Liberman
-
依托单位:
海外基金