课题基金 / 基金详情

Creating a Data Quality Control Framework for Producing New Personnel-Based S&E Indicators

Creating a Data Quality Control Framework for Producing New Personnel-Based S&E Indicators
创建数据质量控制框架以产生新的基于人员的S
批准号:
1917663
负责人:
Jason Owen-Smith
金额:
$46.2万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-09-01 至 2022-08-31

项目摘要

项目成果

Jason Owen-Smith的其他基金

相似基金

相关文献

中文摘要
翻译
科学与工程(S)研究在人类知识、社会效益和经济效益方面产生了可观的回报。全球各国通过大量研究资金和集中努力培养训练有素的劳动力,争夺科学和技术领导地位。迄今为止,衡量和理解S会计师事务所国内和国际趋势以及评估全球优势和劣势的努力,在很大程度上依赖于使用不断增长的大型数据集对专利和出版物等文件进行分析。但这种方法经常遗漏或错误地识别从事富有成效的科学和工程工作的人员和团队。在分析和报告中,很大程度上没有关于S员工队伍规模、构成、协作和跨国流动的强劲指标。侧重于文件和引文的数据分析很难捕捉到国家和国际科学事业的这些关键方面。为了解决这一问题,该项目开发了人员层面的员工队伍和协作措施,可以为S公司的国际竞争力比较增加粒度,并为S公司员工队伍的培训、招聘和留住一个国家的未来带来新的政策见解。这种人员级别指标的先决条件是正确识别和链接出现在多个书目数据集中的个别研究人员。根据作者的名字有效地识别和联系作者是令人望而生畏的,因为名字往往是模棱两可的。亚洲人的名字尤其如此,这带来了一个重大问题,因为亚洲研究人员在许多研究领域发挥着越来越重要的作用。该项目解决了使用新的自动化和层次化实体歧义消除框架在大型书目数据集中系统和例行地消除名称歧义的挑战。这项工作的核心数据集是使用一种新方法得出的,该方法依赖于多个数据字段和迭代过程,以自动创建消除歧义的数据集,这些数据集可用于培训人工智能工具,以进行稳健的人员级别分析。为了提高歧义消除的准确性,根据姓名种族将姓名实例分层为两组,并分别进行歧义消除,以产生基于自动生成的真理数据学习的最佳模型。基于消除歧义的数据,该项目开发了新的人级S指标,该指标表征了所有科学和工程领域的国际S研究队伍的格局和趋势。用于大规模自动消除歧义的新大数据工具将被记录并公开发布,以便于科学界和科学政策研究人员进行扩展、验证和重用。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Science and Engineering (S&E) research generates substantial returns in terms of human knowledge, social and economic benefits. Nations around the globe compete for scientific and technological leadership through substantial research funding and focused efforts to develop highly trained workforces. To date, efforts to measure and understand national and international trends in S&E and to assess global strengths and weaknesses, have largely relied on the analysis of documents such as patents and publications using big, growing datasets. But this approach too often misses or mistakenly identifies the people and teams who do productive science and engineering work. Robust indicators of the size, composition, collaboration, and mobility of the S&E workforce within and across nations are largely missing from analysis and reporting. These key aspects of the national and international scientific enterprise are poorly captured by data analysis focused on documents and citations. To address this problem, this project develops person level workforce and collaboration measures that could add granularity to comparisons of international S&E competitiveness and lead to new policy insights for S&E workforce training, hiring, and retention for a nation's future. The prerequisite of such person level indicators is that individual researchers who appear in multiple bibliographic datasets are correctly identified and linked. Effective identification and linkage of authors based on their names is daunting because names are often ambiguous. This is particularly the case for Asian names, which poses a significant problem as Asian researchers play an increasingly important role in many fields of research. This project addresses the challenge of systematically and routinely disambiguating names in big bibliographic datasets using a new Automated and Stratified Entity Disambiguation framework. Core datasets for this effort are derived using a new method that relies on multiple data fields and an iterative process to automatically create disambiguated datasets that can be used to train artificial intelligence tools to conduct robust person level analysis. To improve disambiguation accuracy, name instances are stratified into two groups according to name-ethnicity and disambiguated separately to produce optimal models learned on the automatically generated truth data. Based on the disambiguated data, this project develops new person-level S&E indicators that characterize the landscape and trends of the international S&E research workforce across all science and engineering fields. The new big data tools for automatic disambiguation at scale will be documented and released publicly to enable expansion, validation, and reuse by the science community as well as science of science policy researchers.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/access.2020.3031112
发表时间: 2020-10
期刊: IEEE Access
影响因子: 3.9
作者: [Jinseok Kim;Jason Owen-Smith]
通讯作者: Jinseok Kim;Jason Owen-Smith
DOI: 10.1177/01655515211018171
发表时间: 2021-05
期刊: Journal of Information Science
影响因子: 2.4
作者: [Jinseok Kim;Jenna Kim;Jinmo Kim]
通讯作者: Jinseok Kim;Jenna Kim;Jinmo Kim
DOI: 10.1007/s11192-020-03826-6
发表时间: 2021-02-11
期刊: SCIENTOMETRICS
影响因子: 3.9
作者: [Kim, Jinseok, Owen-Smith, Jason]
通讯作者: Owen-Smith, Jason
Collaborative Research: RUI: HNDS-R: Stepping out of flatland: Complex networks, topological data analysis, and the progress of science
Collaborative Research: Industries of Ideas: A prototype system for measuring the effects of research investments on regional firms and jobs
ECR: BCSER: IRM: Building Big Data Capacity for Education and Social Science Research Communities Using Restricted Administrative Data
Collaborative Research: Impacts of Hard/Soft Skills on STEM Workforce Trajectories
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    40万元
  • 批准年份:
    2020
  • 负责人:
    Vikrant Gupta
  • 依托单位:
基于Linked Open Data的Web服务语义互操作关键技术
  • 批准号:
    61373035
  • 项目类别:
    面上项目
  • 资助金额:
    77.0万元
  • 批准年份:
    2013
  • 负责人:
    冯志勇
  • 依托单位: