E-MELD: Electronic Metastructure for Endangered Languages Data
E-MELD: Electronic Metastructure for Endangered Languages Data
批准号:
0729644
负责人:
Anthony Aristar
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-03-15 至 2008-06-30
中文摘要
语言数据是一个庞大的社会科学团体研究的核心--不仅是语言学家,而且还有对土著人民文化感兴趣的人类学家、考古学家、历史学家和社会学家。这一研究界的成员目前面临着两个紧迫的情况:世界上的语言数量正在迅速减少,而建立语言数据数字档案的倡议却在迅速增加。面对前者,后者可能看起来是一件纯粹的好事,但如果档案保管员、语言学家和语言工程师之间没有足够的合作,事情可能会出现两种情况。首先,语言数据数字化的共同标准可能永远不会达成一致。由此产生的存档做法和语言表达的差异将严重阻碍数据访问、搜索和跨语言比较。其次,标准的实施可能没有最了解人类语言结构可能性范围的人的指导--对描述不佳的语言做过实地工作的描述性语言学家。如果语言数据和文件的数字档案要提供尽可能广泛的访问并以最有用的形式提供信息,就必须就档案基础设施的某些方面达成共识。作为世界上最大的语言组织和该学科的中央电子出版物,Languist List http://www.linguistlist.org正在组织一个具有双重目标的合作项目:(1)保存濒危语言数据和文件;(2)帮助发展语言档案的基础设施。该项目的一个成果将是建立一个语言学家名单数字档案馆,其中存储着来自10种濒危语言的数据。但对基础设施的关注将产生其他同样重要的结果。首先,语言学家档案将发挥作用,不仅是一个储存库,也是一个最佳实践的陈列室。档案馆将提供根据社区对最佳做法的共识进行标记和编目的濒危语言数据;此外,档案馆还将传播描述最佳做法的参考材料和支持最佳做法的软件工具。另一个成果将是在语言学家名单网站上为该学科设立一个中央元数据服务器;该服务器将整理分布在不同地点的所有与语文有关的资源的信息,而不仅仅是濒危语文的信息。与基础设施有关的其他成果包括:(1)语言学界参与建立最佳做法;(2)广泛传播由此产生的建议;(3)对大量核心语言学家和语文档案保管员进行执行准则方面的实际培训。虽然数据收集工作最初将侧重于濒危语言,但元数据服务器、最佳做法建议以及辅助软件的分发将对语言学的所有实证研究产生重大影响。因此,该项目将为目前计划或正在进行的许多其他与语文有关的项目增值。
英文摘要
Language data is central to the research of a large social sciences community - not only linguists, but also anthropologists, archaeologists, historians, and sociologists interested in the culture of indigenous peoples. Members of this research community are currently faced with two urgent situations: the number of languages in the world is rapidly diminishing while the number of initiatives to create digital archives of language data is rapidly multiplying. The latter might seem to be an unalloyed good in the face of the former, but there are two ways things may go wrong without adequate collaboration among archivists, linguists, and language engineers. First, a common standard for the digitization of linguistic data may never be agreed upon. And the resulting variation in archiving practices and language representation would seriously inhibit data access, searching, and cross-linguistic comparison. Second, standards may be implemented without guidance from the people who best know the range of structural possibilities in human language-descriptive linguists who have done fieldwork on poorly described languages.If digital archives of language data and documentation are to offer the widest possible access and to provide information in a maximally useful form, consensus must be reached about certain aspects of archive infrastructure. As the largest linguistic organization in the world and the central electronic publication of the discipline, The LINGUIST List http://www.linguistlist.org is organizing a collaborative project with a dual objective: (1) to preserve endangered languages data and documentation and (2) to aid in the development of infrastructure for linguistic archives. One outcome of the project will be a LINGUIST List digital archive housing data from 10 endangered languages. But the focus on infrastructure will produce other, equally important results. In the first place, The LINGUIST archive will function, not only as a repository, but also as a 'showroom of best practice.' The archive will offer endangered languages data marked up and catalogued according to community consensus about best practice; furthermore, the archive will disseminate reference material delineating best practice and software tools supporting it. Another outcome will be the establishment on the LINGUIST List site of a central metadata server for the discipline; this server will organize information on all the language-related resources residing at distributed sites, not just endangered languages information alone. Other infrastructure-related outcomes include (1) the involvement of the linguistics community in establishing best practice, (2) the widespread dissemination of the resulting recommendations, and (3) the hands-on training of a substantial core of linguists and language archivists in the implementation of the guidelines. Although the data collection efforts will focus initially on endangered languages, the metadata server, the recommendations for best practice, and the distribution of supporting software will have a significant impact on all empirical research in linguistics. The project will thus add value to many other language-related projects currently planned or underway.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Endangered Languages Information and Infrastructure Project
-
批准号:0924127
-
项目类别:Standard Grant
-
资助金额:$3.96万
-
财政年份:2009
-
负责人:Anthony Aristar
-
依托单位:
Collaborative Research: Multi-Tree: A Digital Library of Language Relationships
-
批准号:0731530
-
项目类别:Continuing Grant
-
资助金额:$19.33万
-
财政年份:2006
-
负责人:Anthony Aristar
-
依托单位:
DHB: Collaborative Research: LL-Map. Language and Location: A Map Annotation Project
-
批准号:0731531
-
项目类别:Standard Grant
-
资助金额:$7.37万
-
财政年份:2006
-
负责人:Anthony Aristar
-
依托单位:
DHB: Collaborative Research: LL-Map. Language and Location: A Map Annotation Project
-
批准号:0527300
-
项目类别:Standard Grant
-
资助金额:$14.44万
-
财政年份:2006
-
负责人:Anthony Aristar
-
依托单位:
Collaborative Research: Multi-Tree: A Digital Library of Language Relationships
-
批准号:0446482
-
项目类别:Continuing Grant
-
资助金额:$25.0万
-
财政年份:2005
-
负责人:Anthony Aristar
-
依托单位:
E-MELD: Electronic Metastructure for Endangered Languages Data
-
批准号:0094934
-
项目类别:Continuing Grant
-
资助金额:$214.29万
-
财政年份:2001
-
负责人:Anthony Aristar
-
依托单位:
Workshop on Endangered Language Data Preservation: The Need for Standards, June 21-24, 2001, Santa Barbara, California
-
批准号:0091713
-
项目类别:Standard Grant
-
资助金额:$4.02万
-
财政年份:2001
-
负责人:Anthony Aristar
-
依托单位:
海外基金