课题基金 / 基金详情

基于引用扩展框架的科学数据可复用性测度研究

批准号:
72104229
项目类别:
青年科学基金项目(C类)
资助金额:
30.0 万元
负责人:
张丽丽
学科分类:
科技管理与政策
结题年份:
2024
批准年份:
2021
项目状态:
已结题
项目参与者:
张丽丽

项目摘要

结项摘要

相似基金

相关文献

中文摘要
数据的可复用性即数据的再利用潜力和过程,是开放科学数据的重要方面。现有的数据可复用性测度缺少低人工干预、人机可读、批量自动化的深度计量方法。数据引用搭建了基于“事实”的“数据-文献”复用链条,但单一指标性能受限,因此需完善测度方法,提升引用对数据可复用性的表征能力。.本研究深入数据引用内容层面,利用主题建模和关键词挖掘,探究语义维度的数据引用位置和功能;采取文献计量学方法,度量时间维度的数据复用演进时效;引入社会关系网络理论,追踪空间维度的数据传播效果。最终建立包含数据引用强度、引用热度、引用广度的三维框架及其三级测度体系。面向数据集、数据库、数据论文等引用场景,完成数据引用扩展框架的基准测试和应用测度。通过对科技文献“致谢”、“数据可用性声明”等短文本挖掘与情感分析,归纳非正式数据引用特征。最终,凝练四种引用场景驱动的数据复用路径与基本规律,形成提升数据可复用性的综合治理方案。
英文摘要
As an important issue for open scientific data, data reusability refers to the capability and process of data reuse. The current data reusability measurements need better-improved data metrics featuring low manual intervention, human-machine readability, and batch automation. Data citation builds a "data-document" reuse chain based on "facts", but the performance of a single indicator is limited. Therefore, it is necessary to enhance the ability of data citation to characterize data reusability...Going deep into the content-level data citation, this research uses topic modeling and keyword mining to explore the location and function of data citation in the semantic dimension; adopts bibliometrics methods to measure the timeliness of data reuse and evolution in the time dimension; introduces social relationship network theory to track the effect of data dissemination in spatial dimensions. Finally, we establish a three-dimensional framework and a three-layer measurement system, including data citation strength, citation heat, and citation breadth. And the benchmark test has been carried out within data citation scenarios, such as data citation for data sets, databases, data papers, and others. Further, the method is employed within the data citation scenarios above. Besides, we use short text mining and sentiment analysis for the "acknowledgment" and "data availability statement" parts within selected research papers to capture the characteristics of informal data citation. In the end, we try to find the paths and basic rules behind data reusability driven by the four data citation scenarios and provide sound suggestions for the sake of better data reusability.
数据可复用性测度是追踪数据开放共享情况的重要工具。基于引用事实链条的数据可复用性计量框架,为更好地表征数据多维度开放复用情况提供了测度方法。本项目重点围绕数据可复用性表征、数据引用扩展框架建模、测试与测度、文本语义挖掘与综合治理方案等核心主题开展研究。. 首先,面向精准计量的数据引用理论研究方面,构建数据引用基本框架,细化引用数据库、引用数据论文、引用数据集、跨库交叉引用等四类数据引用场景,识别五类正式数据引用核心要素以及非正式数据使用情况,聚焦数据粒度、数据流、语义关联与开放引用等,分析数据引用创新技术进展。 . 其次,开展数据引用特征刻画与数据可复用性规律探索。应用文献计量学方法,构建由期刊平台、数据论文、数据集、引证关系等维度在内的数据集、数据论文可复用性计量框架。遴选案例,捕捉数据可复用性特征。引入机器学习方法,设计引文文本自动化采集技术方案,共享语料数据集,凝练形成引文自动分类方法。设计形成由引用位置、引用强度、引用功能、引用情感以及引用时间和空间等指标在内的计量框架,揭示数据可复用性真实情况与发展潜力。开展基于数据特征的数据开放度计量方法研究,构建由科学、技术管理、社会人文、经济、法律伦理等在内的计量体系。. 此外,集成数据资源层、数据平台服务层、数据基础设施生态系统层治理经验,构建面向数据可复用性提升的生态治理协同方案,并创新性地提出减灾场景数据治理的PROTECT原则。. 上述研究,扩展了文献计量学方法在数据可复用性测度方面的应用。重点考量动态数据引用、单篇数据引用、数据引用语义以及数据引用位置与动机,形成由引用功能和引用感情校正的引用强度以及引用热度、引用广度等量化指标,构建形成数据可复用性计量框架,研究形成引文文本自动分类方法,共享引文语料数据集,总结减灾数据治理经验等,为更好地了解、追踪开放数据共享真实情况与发展潜力,提供了积极探索。
国内基金
海外基金