基于学术文献全文内容的细粒度算法实体抽取与评估研究
批准号:
72074113
项目类别:
面上项目
资助金额:
50.0 万元
负责人:
章成志
依托单位:
学科分类:
数字治理与信息资源管理
结题年份:
2024
批准年份:
2020
项目状态:
已结题
项目参与者:
章成志
中文摘要
在大数据时代,算法在学术研究中的提出、改进和应用,对推动不同学科的发展起到重要作用。在基于数据驱动型的研究成果中包含大量的算法,为研究特定领域算法实体提供重要的数据基础。本研究以自然语言处理领域为例,利用该领域学术文献全文内容,开展特定领域算法实体抽取和评估研究。具体而言,首先,本项目构建规范的全文语料和训练语料;随后,利用弱监督机器学习方法从文献中自动抽取算法实体和其他相关实体,并识别算法实体关系、不同实体间关系,在减少人工标注量的同时,保证结果的准确性;接着,统计算法实体被提及频次,识别算法提及位置与动机、算法所在文献的主题和任务;最后,本研究综合上述特征,对算法实体进行多维度关联分析,并构建算法知识库,从而对特定领域的算法进行综合评估。本研究的结果,能为相关科研人员提供科学、客观的视角,帮助他们全面了解不同算法在特定领域的影响力与应用情况,从而为他们选择与使用算法提供坚实的基础。
英文摘要
In the era of big data, the proposal, improvement and application of algorithms in academic research have played an important role in promoting the development of different disciplines. Academic documents in the field of data-driven research contain a large number of algorithms,which are important basic data for researching algorithm entities in specific fields. To this end, this article takes the field of natural language processing (NLP) as an example, and uses the full text of academic papers in NLP to carry out the extraction and evaluation of algorithm entities in a specific field. Specifically, this research manually constructs a structured full-text corpus and training corpus based on content analysis. Then, we use weakly supervised machine learning methods to automatically extract algorithm entities and other entities from academic papers, and explore the relationship between different entities. The method reduce the workload of manual annotation while ensuring the accuracy of the results. At the same time, this research utilizes the full-text content to count the mentioned frequency, identify the location and motivation of the algorithm entity, as well as the topic and task of the article where the algorithm extracted. Finally, based on the above characteristics, multi-dimensional association analysis is performed on algorithm entities, and an algorithm knowledge base is constructed to comprehensively evaluate the algorithm in a specific domain. The results of this research can provide researchers with a scientific and objective perspective and help them understand the influence and practical application of different algorithms in specific fields, thereby providing them with a solid foundation for understanding, selecting, and applying algorithms.
在大数据时代,算法的提出、改进和应用对推动各学科的发展起着关键作用。随着数据驱动型研究成果的增加,算法已成为学术研究的重要组成部分,既增强了科研人员对算法的依赖,也为特定领域算法实体的研究提供了基础。然而,信息过载问题日益严重,传统的依赖人工阅读文献获取算法的方法效率低下。自动识别学术文献中的算法、明确其提及动机并评估影响力,成为亟待解决的难题。. 本项目基于语料库构建,重点研究细粒度算法实体抽取与影响力评估。项目的主要创新点包括:首先,在算法实体自动抽取方面,借助预训练和大模型技术,取得了在多个数据集上的优异表现。同时,项目贡献了多份高质量标注数据,涵盖自然语言处理和图书情报领域的细粒度知识实体标注数据集,以及多模态关键词数据集。. 其次,项目在算法影响力评估层面深入探讨了算法群体的影响力特征,分析了不同时期各算法的作用与影响差异,并从算法提出行为的角度评估各国学术影响力。通过构建算法被提及动机的分类框架,为算法影响力分析提供了新的视角。此外,项目还分析了算法在特定学科领域的扩散趋势,构建了学科研究方法体系,提出了高效的论文推荐和新兴主题识别方法。基于细粒度知识实体和文档多维信息嵌入的算法,相较于现有方法提升了6.7%。. 在篇章结构方面,项目识别了结构功能的关键特征和有效模型,发现基于ChatGPT的数据增强方法能显著提升研究流程段落的识别准确率。此外,项目在具体应用中将新颖性预测、摘要生成等任务与篇章结构有效结合。. 最后,在文本内容学术评价方面,项目探讨了词嵌入模型在新颖性测量中的差异性,发现SciBERT与审稿人新颖性评分的相关性最高。基于这些研究,本项目创新了算法实体的评估方法,通过多因素综合评估,丰富并扩展了现有的评估理论体系。. 此外,基于算法实体抽取,项目构建了细粒度算法知识库,并提供了在线浏览与检索功能,能够为科研人员提供算法在特定领域的影响力与应用情况的科学、客观视角。结合文献主题和任务,项目还帮助确定了常用和最佳算法,从而使学术影响力评估结果更加实用和可靠。
基于可比语料的多语言文本聚类研究
-
批准号:70903032
-
项目类别:青年科学基金项目
-
资助金额:19.0万元
-
批准年份:2009
-
负责人:章成志
-
依托单位:
国内基金
海外基金