Collaborative Research: Implementing the GOLD Community of Practice: Laying the Foundations for a Linguistics Cyberinfrastructure
Collaborative Research: Implementing the GOLD Community of Practice: Laying the Foundations for a Linguistics Cyberinfrastructure
批准号:
0720122
负责人:
Helen Aristar-Dry
金额:
$8.71万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-09-01 至 2011-08-31
中文摘要
语言学科学的经验成分已经看到了数字形式的数据量的快速增长。尽管最近在标记语言、Web协议和数据管理技术方面取得了进展,但语言学作为一个整体还不能充分利用它们。例如,单独的语言数据集通常封装在与其他数据不兼容的形式中:语言数据通常是不可互操作的。这在一定程度上是因为语言学才刚刚开始开发用于管理其数据的领域范围内的最佳实践资源,包括通用软件工具、Web基础设施和知识组件(如本体)。事实上,这些资源将成为任何全领域网络基础设施工作的支柱。为了实现这一目标,这个合作项目将实现GOLD实践社区,这是一个将在线语言数据与语言描述通用本体(GOLD)捕获的语言知识链接起来的Web架构。这个项目的组成部分,由夏飞和威廉·d·刘易斯领导,将通过从网络上收集大量的行间注释文本来解决遗留数据的问题。研究结果将转化为最佳做法格式,并储存在联机行间文本数据库(ODIN)中。其次,Helen Aristar- dry和Anthony Aristar将通过进一步开发FIELD(一种允许现场语言学家生成高质量词汇数据的工具),专注于直接创建最佳实践数据。最后,Scott Farrar将在GOLD框架中实例化得到的最佳实践数据,从而集成来自前两个组件的数据。然后,研究团队将通过实现一个本体驱动的搜索工具来证明项目的有效性,该工具将语言学的一般知识与数据实例捕获的特定知识结合起来。为了确保最终的架构得到广泛的曝光,我们将把这个项目的结果放在LINGUIST List上,在那里它可以被整个语言学社区看到和评估。这个项目将允许普通语言学家和任何对人类语言感兴趣的人在大量的语言数据中搜索和查看概括。它将直接解决涉及比较和集成数据的关键问题,这些问题最初并不打算进行比较。这些包括利用现有资源(例如,来自Web的遗留数据)、利用最佳实践数据标准以及利用领域范围的知识。这些问题提出了重大的技术挑战,因为对于给定的领域(如语言学)没有通用的现成解决方案。项目的成功需要对语言数据对象和结构有深刻的理解。事实上,该项目将展示如何在更广泛的框架中利用该领域的基本数据结构。在这个世界即将失去语言多样性的时代,这个项目将产生一个社区范围内可用的资源,作为探索各种人类语言结构的搜索工具。目前,还没有这样的语言数据搜索工具。当用户看到为这样的工作做出贡献的价值时,他们将更有可能接受附带的数据标准和工具。因此,该项目将实现的是一个致力于生产高质量数据资源的语言学家社区,其共同目标是影响我们对语言结构的理解的下一个巨大进步。
英文摘要
The empirical component of the linguistics sciences has seen a rapid increase in the amount of data available in digital form. Though there have been recent advances in markup languages, Web protocols, and techniques for data management, linguistics as a whole has not been able to take full advantage of them. For instance, individual sets of linguistic data are often encapsulated in forms that are not compatible with others: linguistic data are not generally interoperable. This is in part because linguistics has only begun to develop field-wide, best-practice resources for managing its data, including common software tools, Web infrastructures, and knowledge components such as ontologies. Such resources would, in fact, act as the backbone for any field-wide cyberinfrastructure effort. Towards such a goal, then, this collaborative project will implement the GOLD Community of Practice, a Web architecture for linking on-line linguistic data to linguistic knowledge captured by the General Ontology for Linguistic Description (GOLD). The component of the project, led by Fei Xia and William D. Lewis, will address the issue of legacy data by harvesting large amounts of interlinear glossed text from the Web. The results will be transformed into a best-practice format and stored in the Online Database of INterlinear text (ODIN). Second, Helen Aristar-Dry and Anthony Aristar will focus on the direct creation of best-practice data by further developing FIELD, a tool that allows field linguists to produce high quality lexical data. Finally, Scott Farrar will instantiate the resulting best-practice data in the GOLD framework, thus integrating data from the first two components. The research team will then demonstrate the efficacy of project by implementing an ontology-driven search facility that incorporates the general knowledge of linguistics with the specific knowledge captured by the data instances. To ensure that the resulting architecture gets wide exposure, we will house the results of this project at the LINGUIST List where it can be seen and evaluated by the linguistics community as a whole.This project will allow ordinary working linguistics and anyone with an interest in human language to search and see generalizations across large amounts of linguistic data. It will directly address the key issues involved in the comparison and integration data that were not originally intended to be comparable. These include the leveraging of existing resources (i.e., legacy data from the Web), taking advantage of best-practice data standards, and utilizing field-wide knowledge. These issues present significant technological challenges, as there are no general off-the-shelf solutions for given domains such as linguistics. The success of the project requires a deep understanding of linguistic data objects and structures. In fact, the project will demonstrate how the fundamental data structures of the field can be utilized in a broader framework. At a time when the world stands to lose much of its linguistic diversity, this project will result in a community-wide resource usable for its intrinsic value as a search tool to explore the structure of all kinds of human languages. At the present, there are no such search tools available for linguistic data. When users see the value in contributing to such an effort, they will be more likely to embrace the accompanying data standards and tools. Thus, what the project will achieve is a community of linguists dedicated to the production of quality data resources for the common goal of affecting the next great advance in our understanding of the structures of language.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
ICE (Integrating Cartographic Elements: Creating Resources Emphasizing Arctic Materials)
-
批准号:0952335
-
项目类别:Standard Grant
-
资助金额:$32.29万
-
财政年份:2009
-
负责人:Helen Aristar-Dry
-
依托单位:
INTEROP: Lexicon Enhancement via the GOLD Ontology (LEGO)
-
批准号:0753321
-
项目类别:Continuing Grant
-
资助金额:$63.64万
-
财政年份:2008
-
负责人:Helen Aristar-Dry
-
依托单位:
Collaborative Research: Workshop: Towards the Interoperability of Language Resources
-
批准号:0709680
-
项目类别:Standard Grant
-
资助金额:$1.33万
-
财政年份:2007
-
负责人:Helen Aristar-Dry
-
依托单位:
DHB: Collaborative Research: LL-Map. Language and Location: A Map Annotation Project
-
批准号:0527512
-
项目类别:Standard Grant
-
资助金额:$59.82万
-
财政年份:2006
-
负责人:Helen Aristar-Dry
-
依托单位:
Collaborative Research: Multi-Tree: A Digital Library of Language Relationships
-
批准号:0445714
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Helen Aristar-Dry
-
依托单位:
DATA: Dena'ina Archiving, Training, and Access
-
批准号:0326805
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:Helen Aristar-Dry
-
依托单位:
Collaborative Project: The Rosetta Project- ALL Language Archive
-
批准号:0333530
-
项目类别:Continuing Grant
-
资助金额:$9.65万
-
财政年份:2003
-
负责人:Helen Aristar-Dry
-
依托单位:
SGER: Database Design for Endangered Languages Data
-
批准号:0003197
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2000
-
负责人:Helen Aristar-Dry
-
依托单位:
The LINGUIST Multi-List Support Project
-
批准号:9975299
-
项目类别:Standard Grant
-
资助金额:$16.8万
-
财政年份:1999
-
负责人:Helen Aristar-Dry
-
依托单位:
Software Development for the LINGUIST Network
-
批准号:9601352
-
项目类别:Standard Grant
-
资助金额:$11.5万
-
财政年份:1996
-
负责人:Helen Aristar-Dry
-
依托单位:
LINGUIST Software Development
-
批准号:9311748
-
项目类别:Standard Grant
-
资助金额:$0.4万
-
财政年份:1993
-
负责人:Helen Aristar-Dry
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: