面向App软件开发的API通用表示模型预训练方法及其应用研究
批准号:
62102160
项目类别:
青年科学基金项目(C类)
资助金额:
30.0 万元
负责人:
刘宇舟
依托单位:
学科分类:
软件理论、软件工程与服务
结题年份:
2024
批准年份:
2021
项目状态:
已结题
项目参与者:
刘宇舟
中文摘要
API的使用是提高App产品开发效率、提升产品质量的重要途径。如何合理使用API资源一直以来都是软件工程领域重要的研究课题。研究表明,利用机器学习方法从大规模数据中获取API使用知识是解决该问题的有效途径。近年来,BERT、GPT-3等模型的出现引发深度学习模式的新浪潮,通过超大规模的模型预训练,能够实现针对不同问题,以“低花费”获取“高质量”的可用模型。本项目拟基于App商店中超大规模代码数据,研究API通用表示模型的预训练方法,利用App特征信息引导预训练过程,克服代码数据训练中的难点问题。项目预期实现两个目标:1、提供通用的API表示预训练模型,阐明其预训练方法并提供相关数据集;2、针对API推荐,API误用预警以及API迁移分析三个具体任务进行模型微调,在检验API表示模型预训练效果的同时,解决App开发中的实际API使用问题。
英文摘要
The API (Application Programming Interfaces) is important for improving the efficiency and quality of App development, and how to use the APIs properly is always a hot research question in Software Engineering. With the development of machine learning, many studies show that it is an efficient way for solving this problem. Recently, BERT and GPT-3 trigger a new wave of deep learning models: after the large-scale pre-training of the model, it can be fine tuned with just one additional output layer to create state-of-the-art models for different tasks, which means we can spend lower cost to gain higher quality models. Following this idea, this project aims to pre-training a general representation model of APIs based on the large-scale of data in App stores. Considering the Apps take the development of “features” as the core, we want to introduce the information of App features to tackle the difficulties caused by the characteristics of codes in the training process (such as code divisions, input embedding, and pre-training task), so that we can gain the pre-trained models with lower cost. The project is intended to achieve two goals: 1. providing a usable pre-trained model of API presentations, and illustrating the training method as well as the training data; 2. fine-tuning the model for three specific tasks, including API recommendation, API misuse detection and API migration analysis, in this way, we could not only validate the performance of the pre-trained model, but also help developers solving the practice problems of using API resource in the development of Apps.
本项目以App商店中大规模数据为基础,研究API通用表示模型的预训练方法,并结合Android软件开发过程中API使用问题设计模型的微调应用方式。为此,我们首先探索了App产品自然语言描述文本和实现代码两种模态中API功能相关的语义信息表示,构建了以产品特征为索引的层次化API知识体系;其次,进一步探究了API结构信息作用,将使用API实现功能这一任务认为视为多目标优化问题,即找到与待实现功能语义和结构都相关的API知识,进而使用遗传算法来获得最优解,实现小规模数据下API语义与结构信息的分析表示;最后,提出构建双模态深度玻尔兹曼机模型,从多模态融合的角度获取更加全面地API通用表示,并在API推荐,相似API挖掘等多种API任务上实现了模型的微调应用。此外,本项目提出了基于敏感API多模态信息的程序恶意行为检测方法,将API表示与程序执行码中信息相结合,不仅提高了恶意行为检测性能,也扩宽了通用模型的应用领域。基于上述研究,项目在国际知名期刊上共发表论文8篇,授权国家发明专利1项。项目的实施完善了现有软件智能化开发领域的理论与方法体系,提出的通用模型及恶意软件检测技术可以为实际软件开发提供技术支持,且有望将技术推广到网络安全等相关领域。
国内基金
海外基金