课题基金 / 基金详情

Developing Standards for Data Citation and Attribution for Reproducible Research in Linguistics

Developing Standards for Data Citation and Attribution for Reproducible Research in Linguistics
为语言学研究的可重复性研究制定数据引用和归因标准
批准号:
1447886
负责人:
Andrea Berez-Kroeker
金额:
$9.84万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-11-15 至 2020-04-30

项目摘要

项目成果

Andrea Berez-Kroeker的其他基金

相似基金

相关文献

中文摘要
翻译
该项目支持一系列的三个研讨会和一个小组报告,将相关利益相关者聚集在一起,制定和促进语言数据的数据引用和归因标准。语言学是一门数据驱动的社会科学,它从对语言实践的观察中得出关于人类认知和社会结构的推论。这些观察,以记录和相关注释的形式,代表了基础字段的主要数据集。这种做法源于文献学,它依赖文本作为主要的数据来源。然而,三个最近且相互关联的因素使语言学的数据导向模型在当前的领域特别相关。首先,技术的重大转变导致了数字语言数据量的迅速增长。第二,超过一半的世界?美国的语言正濒临灭绝,因此在不久的将来,档案数据将是这些语言信息的唯一来源。第三,文献语言学作为一个公认的子领域的出现导致了对数据管理和管理的更多关注。虽然语言学家一直依赖于语言数据,但他们并不总是为获取这些数据提供便利。语言学出版物通常包括数据集的简短摘录,通常由少于五个单词组成,并且通常没有引用。在提供引文的地方,通常只能模糊地识别与数据集的联系。一段摘录可能会有一个引用,引用的是摘录的文本的名称,但实际上读者没有办法访问该文本。也就是说,尽管该领域最近的变化产生了潜力,但今天创造的大量语言学研究无论是在原则上还是在实践中都是不可复制的。研讨会和小组报告将促进语言学数据管理和引用标准的发展,以响应这些不断变化的条件,并将语言学领域转向更科学,数据驱动的模型,从而产生可重复的研究。阻碍语言学可重复研究发展的一个主要因素是缺乏数据引用和归因标准。尽管人们越来越认识到语言数据的重要性,但对这些数据的引用还没有广泛建立的指导方针。同样重要的是,没有归因标准。缺乏这样的标准,期刊、学术终身职位和晋升委员会以及同行评审过程继续强调语言分析而不是语言数据,因此语言学家没有动力使数据易于获取。数据驱动的语言科学有可能通过促进对语言数据的关注和结构来为科学主张提供证据。到项目结束时,研究人员将举行三次研讨会,研究和开发语言学中数据引用和归因的模型;在美国语言学会2017年年会上促进了关于这些主题的学科范围的讨论;撰写关于语言学引文和归因标准的意见书;并提交了一份关于引用和归属于LSA的决议提案。
英文摘要
This project supports a series of three workshops and one panel presentation bringing together relevant stakeholders to develop and promote standards for data citation and attribution for linguistic data. Linguistics is a data-driven social science, in which inferences about human cognition and social structure are drawn from observations of linguistic practice. These observations, in the form of recordings and associated annotations, represent the primary data sets that underlie the field. This practice has its roots in philology, which relies on texts as a primary data source. However, three recent and inter-related factors make the data-oriented model of linguistics particularly relevant to the field at the current time. First, a major shift in technology has resulted in rapidly growing volumes of digital language data. Second, more than half of the world?s languages are critically endangered, so that in the not-so-distant future archival data will be the only source of information on those languages. Third, the emergence of Documentary Linguistics as a recognized sub-field has led to an increased focus on data curation and management. While linguists have always relied on language data, they have not always facilitated access to those data. Linguistic publications typically include short excerpts from data sets, ordinarily consisting of fewer than five words, and often without citation. Where citations are provided, the connection to the data set is usually only vaguely identified. An excerpt might be given a citation which refers to the name of the text from which it was extracted, but in practice the reader has no way to access that text. That is, in spite of the potential generated by recent shifts in the field, a great deal of linguistic research created today is not reproducible, either in principle or in practice. The workshops and panel presentation will facilitate development of standards for the curation and citation of linguistics data that are responsive to these changeing conditions and shift the field of linguistics toward a more scientific, data-driven model which results in reproducible research. A primary factor hindering the development of reproducible research in linguistics is the lack of standards for data citation and attribution. Although language data are increasingly recognized as important, there are no widely established guidelines for the citation of these data. Equally important, there are no standards for attribution. Lacking such standards, journals, academic tenure and promotion committees, and peer review processes continue to emphasize linguistic analyses over linguistic data, and as a result linguists have little incentive to make data accessible. A data-driven linguistic science has the potential to provide substantiation of scientific claims by promoting attention to the care and structuring of language data. By the end of the project, the researchers will have held three workshops to research and develop a model for data citation and attribution in linguistics; facilitated discipline-wide discussion on these topics at the 2017 annual meeting of the Linguistic Society of America; written a position paper on standards for citation and attribution in linguistics; and submitted a proposal for a Resolution on citation and attribution to the LSA.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1515/ling-2017-0032
发表时间: 2018-01-01
期刊: LINGUISTICS
影响因子: 1.1
作者: [Berez-Kroeker, Andrea L., Gawne, Lauren, Woodbury, Anthony C.]
通讯作者: Woodbury, Anthony C.
Doctoral Dissertation Research: Integration of Quantitative and Documentary Methodologies in the Analysis of a Segmentally-Rich Language
  • 批准号:
    1840668
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.73万
  • 财政年份:
    2019
  • 负责人:
    Andrea Berez-Kroeker
  • 依托单位:
Doctoral Dissertation Research: Child and Child-Directed Expression of Possession in a Polysynthetic Language
  • 批准号:
    1912062
  • 项目类别:
    Standard Grant
  • 资助金额:
    $2.12万
  • 财政年份:
    2019
  • 负责人:
    Andrea Berez-Kroeker
  • 依托单位:
RR: EAGER: Data Science Literacy for All of Linguistics
  • 批准号:
    1745249
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.1万
  • 财政年份:
    2017
  • 负责人:
    Andrea Berez-Kroeker
  • 依托单位:
Vital Voices: Linking Language and Wellbeing at the International Conference on Language Documentation and Conservation
  • 批准号:
    1614134
  • 项目类别:
    Standard Grant
  • 资助金额:
    $6.0万
  • 财政年份:
    2016
  • 负责人:
    Andrea Berez-Kroeker
  • 依托单位:
海外基金