课题基金 / 基金详情

SOCS: Socially Intelligent Computing for Coding of Qualitative Data

SOCS: Socially Intelligent Computing for Coding of Qualitative Data
SOCS:用于定性数据编码的社会智能计算
批准号:
1111107
负责人:
Nancy McCracken
金额:
$74.78万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-09-01 至 2015-08-31

项目摘要

项目成果

Nancy McCracken的其他基金

相似基金

相关文献

中文摘要
翻译
该项目将开发和评估基于自然语言处理(NLP)和机器学习(ML)的创新研究工具,以支持定性社会科学研究,特别是内容分析。内容分析是一种定性研究技术,用于使用文本而不是数字作为其原始数据来寻找理论感兴趣的概念的证据。 识别和标记文本中重要特征的过程称为“编码”,这种分析的结果是用所展示概念的代码注释的文本。这种技术已经变得越来越受欢迎和更适用的数量可用的“数字化”文本已经爆炸。 然而,对文本的手动分析的依赖限制了内容分析研究的规模和范围。在这个项目中,编码定性数据的问题被概念化为一个信息提取问题,适合使用NLP自动化。然而,这些技术将被用于支持角色,而不是寻求自动化过程,从而创建人机合作伙伴关系。ML将用于从编码文本的示例中归纳NLP规则,避免手动开发规则的需要。为了减少人类参与者所需的训练数据量,将采用主动学习过程,其中使用一些手工编码的示例来创建初始模型,该模型可以通过与用户的交互进一步发展。这些方法将结合在一个原型工具中,以支持定性内容分析。作为该工具的演示和测试,它将被应用于网络基础设施支持的分布式群体,特别是自由/自由开源软件开发团队的当前和新的研究,然后到广泛的社会科学研究问题。这种广泛的使用也将提供一个测试的社会计算方法来解决这个问题的普遍性。 首先,该提案旨在开发一种新的社会计算系统,通过整合信息提取和主动学习来支持人机伙伴关系。第二,验证研究将适用于一组不同的代码的工具,提供证据的社会计算方法的一般性和局限性。 第三,使用该工具的示范研究将有助于分布式群体的研究。 最后,该项目解决了一个基本的方法问题,在广泛的领域的定性研究,即处理大量的非结构化的定性数据,通过应用创新的计算机支持。 通过避免手写规则的需要和减少所需的手写注释培训数据量,这一伙伴关系将实际使用一个系统,对各个领域的大量定性数据进行编码。它将通过为科学研究提供内容分析工具形式的研究提供有用的基础设施,并以注释数据语料库的形式用于未来的自然语言处理研究,从而使社会受益。示范研究将提供可推广的知识,以提高分布式群体的有效性,这是一种日益重要的组织模式。最后,该项目有助于教育和培训,特别是对妇女和少数群体成员的教育和培训。
英文摘要
This project will develop and evaluate an innovative research tool, based on Natural Language Processing (NLP) and Machine Learning (ML), to support qualitative social science research, specifically content analysis. Content analysis is a qualitative research technique for finding evidence of concepts of theoretical interest using text rather than numbers as its raw data. The process of identifying and labeling significant features in text is referred to as "coding," and the result of such an analysis is a text annotated with codes for the concepts exhibited. This technique has become increasingly popular and more applicable as the volume of available "born-digital" text has exploded. However, the reliance on manual analysis of the text limits the scale and scope of content analysis research.In this project, the problem of coding qualitative data is conceptualized as an information extraction problem amenable to automation using NLP. However, rather than seeking to automate the process, the technologies will be used in a supporting role, creating a human-computer partnership. ML will be used to induce NLP rules from examples of coded text, avoiding the need to develop rules manually. To reduce the amount of training data needed from the human participants, an active learning process will be employed, in which a few hand-coded examples are used to create an initial model that can be further evolved through interaction with the user. These approaches will be combined in a prototype tool to support qualitative content analysis. As a demonstration and test of the tool, it will be applied to current and novel studies of cyber-infrastructure-supported distributed groups, specifically free/libre open source software development teams, and then to a broad range of social science research problems. This broad usage will also provide a test of the generalizability of a socio-computational approach to this problem.The intellectual merit of the research is four-fold. First, the proposal seeks to develop a novel socio-computational system that supports a human-computer partnership through the integration of information extraction and active learning. Second, a validation study will apply the tool to a diverse set of codes, providing evidence of the generality and limits of a socio-computational approach. Third, the demonstration studies using the tool will contribute to research on distributed groups. Finally, the project addresses a fundamental methodological problem in the broad domain of qualitative research, namely dealing with large quantities of unstructured qualitative data, by applying innovative computer-support. By avoiding the need for hand-written rules and reducing the required amount of hand-annotated training data, this partnership will make practical the use of a system for coding large quantities of qualitative data in various domains.The project has numerous broader impacts. It will benefit society by providing useful infrastructure for research in the form of a content analysis tool for scientific research and in for the form of corpora of annotated data for use in future Natural Language Processing research. The demonstration studies will provide generalizable knowledge to improve the effectiveness of distributed groups, an increasingly important mode of organization. Finally, the project contributes to the education and training, of women and minority group members in particular.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
The Design, Definition, and Implementation of Programming Languages
  • 批准号:
    8604177
  • 项目类别:
    Standard Grant
  • 资助金额:
    $3.1万
  • 财政年份:
    1986
  • 负责人:
    Nancy McCracken
  • 依托单位:
Typechecking and Implementation Properties of Programming Languages With Extended Type Structure
  • 批准号:
    8004219
  • 项目类别:
    Standard Grant
  • 资助金额:
    $4.01万
  • 财政年份:
    1980
  • 负责人:
    Nancy McCracken
  • 依托单位:
海外基金