课题基金 / 基金详情

SOCS: Socially Intelligent Computing for Coding of Qualitative Data

SOCS: Socially Intelligent Computing for Coding of Qualitative Data
SOCS:用于定性数据编码的社会智能计算
批准号:
1111107
负责人:
Nancy McCracken
金额:
$74.78万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-09-01 至 2015-08-31

项目摘要

项目成果

Nancy McCracken的其他基金

相似基金

相关文献

中文摘要
翻译
该项目将开发和评估一个基于自然语言处理(NLP)和机器学习(ML)的创新研究工具,以支持社会科学的定性研究,特别是内容分析。内容分析是一种使用文本而不是数字作为原始数据来寻找理论上感兴趣的概念的证据的定性研究方法。识别和标记文本中的重要特征的过程被称为“编码”,并且这种分析的结果是带有所示概念的代码的注释的文本。随着可用的“天生数字”文本的数量激增,这种技术变得越来越流行,也越来越适用。然而,对文本的人工分析的依赖限制了内容分析研究的规模和范围。在这个项目中,定性数据的编码问题被概念化为一个信息提取问题,可以使用NLP进行自动化。然而,这些技术将作为辅助角色,创建人机合作伙伴关系,而不是寻求实现这一过程的自动化。ML将用于从编码文本的示例中归纳出NLP规则,从而避免了手动开发规则的需要。为了减少从人类参与者那里需要的训练数据量,将采用主动学习过程,其中使用几个手工编码的例子来创建初始模型,该初始模型可以通过与用户的交互来进一步进化。这些方法将结合在一个原型工具中,以支持定性内容分析。作为对该工具的演示和测试,它将应用于对网络基础设施支持的分布式群体的当前和新的研究,特别是自由/libre开放源码软件开发团队,然后应用于广泛的社会科学研究问题。这种广泛的使用也将测试社会计算方法对这个问题的普适性。这项研究的学术价值有四个方面。首先,该提案寻求开发一种新的社会计算系统,通过整合信息提取和主动学习来支持人机伙伴关系。其次,验证性研究将该工具应用于一组不同的代码,提供社会计算方法的普遍性和局限性的证据。第三,使用该工具的示范研究将有助于对分布式组的研究。最后,该项目解决了定性研究广泛领域中的一个基本方法问题,即通过应用创新的计算机支持来处理大量非结构化定性数据。通过避免需要手写规则并减少所需的手写注解训练数据量,这一伙伴关系将实际使用一个系统对不同领域的大量定性数据进行编码。该项目具有许多更广泛的影响。它将为研究提供有用的基础设施,为科学研究提供内容分析工具,并以注释数据语料库的形式为未来的自然语言处理研究提供有用的基础设施。示范研究将提供一般性知识,以提高分布式组织的有效性,分布式组织是一种日益重要的组织模式。最后,该项目有助于教育和培训,特别是对妇女和少数群体成员。
英文摘要
This project will develop and evaluate an innovative research tool, based on Natural Language Processing (NLP) and Machine Learning (ML), to support qualitative social science research, specifically content analysis. Content analysis is a qualitative research technique for finding evidence of concepts of theoretical interest using text rather than numbers as its raw data. The process of identifying and labeling significant features in text is referred to as "coding," and the result of such an analysis is a text annotated with codes for the concepts exhibited. This technique has become increasingly popular and more applicable as the volume of available "born-digital" text has exploded. However, the reliance on manual analysis of the text limits the scale and scope of content analysis research.In this project, the problem of coding qualitative data is conceptualized as an information extraction problem amenable to automation using NLP. However, rather than seeking to automate the process, the technologies will be used in a supporting role, creating a human-computer partnership. ML will be used to induce NLP rules from examples of coded text, avoiding the need to develop rules manually. To reduce the amount of training data needed from the human participants, an active learning process will be employed, in which a few hand-coded examples are used to create an initial model that can be further evolved through interaction with the user. These approaches will be combined in a prototype tool to support qualitative content analysis. As a demonstration and test of the tool, it will be applied to current and novel studies of cyber-infrastructure-supported distributed groups, specifically free/libre open source software development teams, and then to a broad range of social science research problems. This broad usage will also provide a test of the generalizability of a socio-computational approach to this problem.The intellectual merit of the research is four-fold. First, the proposal seeks to develop a novel socio-computational system that supports a human-computer partnership through the integration of information extraction and active learning. Second, a validation study will apply the tool to a diverse set of codes, providing evidence of the generality and limits of a socio-computational approach. Third, the demonstration studies using the tool will contribute to research on distributed groups. Finally, the project addresses a fundamental methodological problem in the broad domain of qualitative research, namely dealing with large quantities of unstructured qualitative data, by applying innovative computer-support. By avoiding the need for hand-written rules and reducing the required amount of hand-annotated training data, this partnership will make practical the use of a system for coding large quantities of qualitative data in various domains.The project has numerous broader impacts. It will benefit society by providing useful infrastructure for research in the form of a content analysis tool for scientific research and in for the form of corpora of annotated data for use in future Natural Language Processing research. The demonstration studies will provide generalizable knowledge to improve the effectiveness of distributed groups, an increasingly important mode of organization. Finally, the project contributes to the education and training, of women and minority group members in particular.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
The Design, Definition, and Implementation of Programming Languages
  • 批准号:
    8604177
  • 项目类别:
    Standard Grant
  • 资助金额:
    $3.1万
  • 财政年份:
    1986
  • 负责人:
    Nancy McCracken
  • 依托单位:
Typechecking and Implementation Properties of Programming Languages With Extended Type Structure
  • 批准号:
    8004219
  • 项目类别:
    Standard Grant
  • 资助金额:
    $4.01万
  • 财政年份:
    1980
  • 负责人:
    Nancy McCracken
  • 依托单位:
海外基金