汉语小句复合体结构自动分析研究
批准号:
62076037
项目类别:
面上项目
资助金额:
58.0 万元
负责人:
罗智勇
依托单位:
学科分类:
自然语言处理
结题年份:
2024
批准年份:
2020
项目状态:
已结题
项目参与者:
罗智勇
中文摘要
汉语从文字、词语、小句到句子等都与印欧语言有很大的不同。在小句和句子层面,汉语缺少话语范畴中小句和句子界定的操作理论,因而更缺少小句和句子自动分析的模型和算法。课题组以往研究表明:汉语句子的切分问题、以及由小句成分共享关系导致的文本远距离依赖问题,仍然是深度学习方法在深层次中文信息处理中面临的困难。本项目拟在小句复合体形式理论的基础上,以标点句为基本单位,以话头话体关系为基本线索,以汉语小句话头共享模式为形式约束,将汉语小句二维邻接特征等远距离相关性规律融入到机器学习模型中,研究汉语句子的边界识别、小句复合体层次结构分析、小句复合体内部话头话体共享关系识别,以及汉语小句动态生成机制。该研究将获得汉语小句复合体结构自动分析的计算模型和算法,使得对于给定的书面文本,机器能自动切分出小句复合体,拆分出小句,以便支持中文信息处理的深入应用。
英文摘要
Chinese is very different from Indo-European language in terms of characters, words, clauses and sentences. At the level of clauses and sentences, Chinese lacks the operational theory of defining clauses and sentences in the discourse category, and thus lacks the models and algorithms for automatic analysis of clauses and sentences. Previous research by the research group shows that the segmentation of Chinese sentences and the long-distance dependence of texts caused by the sharing of clause components are still the difficulties faced by deep.learning methods in in-depth Chinese information processing. Based on the formal theory of clause complexes, this project intends to integrate the long-distance correlation laws such as the 2-dimensional adjacency features of Chinese clauses into the machine learning model, with punctuation clauses as the basic unit, the relation between the Naming and Telling as the basic clue, and the Naming-sharing mode of Chinese clauses as the formal constraint. It will study the boundary identification of Chinese sentences, the hierarchical structure analysis of clause.complexes, the identification of the Naming-sharing relationship within clause complexes, and the dynamic generation of Chinese clauses. This research will obtain the computational model and algorithm for automatic analysis of Chinese clause complex structure, so that for a given written text, the machine can automatically segment the clause complex and the clause, so as to support the in-depth application of Chinese information processing.
汉语从文字、词语、小句到句子等都与印欧语言有很大的不同。在小句和句子层面,汉语缺少话语范畴中小句和句子界定的操作理论,因而更缺少小句和句子自动分析的模型和算法。课题组在小句复合体形式理论的基础上,以标点句为基本单位,以话头话体关系为基本线索,以汉语小句话头共享模式为形式约束,将汉语小句二维邻接特征等远距离相关性规律融入到机器学习模型中,研究汉语句子的边界识别、小句复合体层次结构分析、小句复合体内部话头话体共享关系识别,以及汉语小句动态生成机制,获得多方面原创性成果,包括:创新性的提出了小句复合体话头话体关系形式化表示模型:NTCGraph,将小句复合体中标点句间话头话体共享关系抽象为广义有向无环图结构;设计和实现了中文小句复合体结构自动分析形式模型和算法,将图结构(主要是有向边)预测问题定义为文本片段抽取(text-span extraction)机器学习问题;创新性的提出NT-MASK注意力机制,将NTCGraph结构预测模型与大规模预训练语言模型深度融合;将小句复合体结构自动识别模型应用于机器阅读理解任务,对远距离跨标点句问答问题有明显效果,与基准模型相比,基于小句复合体的机器阅读理解模型的整体精确匹配率(EM)提升了3.26%,其中跨标点句问答问题的EM提升了3.49%。.课题组共出版专著2部,公开发表论文8篇(其中EI索引5篇);培养硕士研究生8名;基于本项目研究成果,达成1项横向技术研究合作项目。研制大规模汉语小句复合体语料库和汉语包孕子句语料库。
国内基金
海外基金