课题基金 / 基金详情

面向数据库自然语言查询的语义理解研究

批准号:
62106142
项目类别:
青年科学基金项目(C类)
资助金额:
30.0 万元
负责人:
陈露
依托单位:
学科分类:
自然语言处理
结题年份:
2024
批准年份:
2021
项目状态:
已结题
项目参与者:
陈露

项目摘要

结项摘要

陈露的其他基金

相似基金

相关文献

中文摘要
近年来发展起来的基于自然语言对话技术的数据库查询方法为用户通过人机对话形式进行数据分析提供了可能。语义理解是其中的核心技术,其作用是将自然语言查询转换为结构化查询语言SQL语句。由于SQL表达形式及数据库结构的复杂多变,当前该语义理解任务面临着三个方面的挑战:标注数据难度大、效率低、成本高;准确的语义理解需要深度利用复杂多变的数据库结构知识;半监督和无监督范式的算法研究不足,缺乏提升训练数据稀疏情况下模型鲁棒性的有效机制。为此,在数据标注方面,本项目拟构建可视化标注平台,提出基于主动学习的数据标注框架,提升标注效率、降低标注难度;在模型结构方面,提出针对异构图的高阶图注意力神经网络,提升模型对异构信息的编码能力;在优化算法方面,提出基于对偶学习的闭环优化框架,提高模型在训练数据稀疏场景下的鲁棒性。通过上述系列研究,本项目力求为构建准确度高、鲁棒性好的语义理解模型提出完整的解决方案。
英文摘要
In recent years, the database query method based on natural language dialogue technology has been developed, which makes it possible for users to analyze data using natural language. Here semantic parsing, i.e. Text2SQL, plays a core role, and its role is to convert natural language query into the structured query language SQL. Due to the complexity and changeability of SQL expression form and database structure, the semantic parsing task is faced with three challenges: 1) Labelling large-scale training data is difficult, low efficiency, and high cost. 2) How to use the complex database structure knowledge to assist natural language understanding is a question. 3) The research of semi-supervised and unsupervised learning algorisms is insufficient, and there is no effective mechanism to improve the model robustness under sparse training data. Therefore, 1) for data annotation, this project plans to build a web-based annotation platform and propose a data annotation framework based on an active learning framework to improve annotation efficiency and reduce annotation difficulty. 2) For model structure, a high-order graph attention neural network for heterogeneous graphs is proposed to improve the ability to encode heterogeneous information. 3) For optimization algorithms, a closed-loop optimization framework based on dual learning is proposed to improve the robustness of the model when training data is sparse. Through the above series of studies, this project tries to put forward a complete solution for constructing a semantic parsing model with high accuracy and good robustness.
近年来发展起来的基于自然语言对话技术的数据库查询方法为用户通过人机对话形式进行数据分析提供了可能。语义理解是其中的核心技术,其作用是将自然语言查询转换为结构化查询语言SQL语句,即Text2SQL。本项目以构建准确度高、鲁棒性好的Text2SQL语义理解模型为目标,围绕数据库和SQL结构复杂、高质量标注数据稀疏等核心挑战,提出了线图增强的对偶注意力图神经网络模型和语法规则引导的SQL语句动态自适应解码算法,提升了模型对复杂输入输出模式结构的建模能力;首次定义了Text2SQL任务的数据库模式结构的鲁棒性问题,发现了影响模型鲁棒性的原因,构建并开源了首个针对数据库模式结构鲁棒性的Text2SQL数据集;探索了基于大模型的Text2SQL研究新范式,提出了基于思维链和编辑链的Text2SQL语义解析新方法,降低了模型对高质量标注数据的需求,提升了语义解析的准确性和推理过程的可解释性。基于上述研究成果,本项目在TPAMI、ICML、NeurIPS、ACL等重要国际期刊和会议以及《中国科学基金》等国内期刊上发表论文12篇,申请国家发明专利4项。项目组所提出的方法和模型在CSpider、DuSQL、NL2SQL、CoSQL等国内外Text2SQL语义解析权威测试基准榜单上都曾取得了最好成绩,并将模型和代码在Github上开源,多个开源成果受到了Text2SQL研究社区的广泛关注。
METTL3-ORC6-CDC45轴介导的前列腺癌干性增强及恩杂鲁胺耐药的机制研究
  • 批准号:
    --
  • 项目类别:
    面上项目
  • 资助金额:
    52万元
  • 批准年份:
    2022
  • 负责人:
    陈露
  • 依托单位:
SIRT5-DBT功能调控轴通过AR-V7参与前列腺癌去势抵抗的机制研究
  • 批准号:
    82072844
  • 项目类别:
    面上项目
  • 资助金额:
    55.0万元
  • 批准年份:
    2020
  • 负责人:
    陈露
  • 依托单位:
国内基金
海外基金