课题基金 / 基金详情

A Lebesgue Integral based Approximation for Language Modelling

A Lebesgue Integral based Approximation for Language Modelling
基于勒贝格积分的语言建模近似
批准号:
EP/X019063/1
负责人:
Lin Gui
金额:
$25.77万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
近年来,基于深度学习的自然语言处理技术引起了人们的极大兴趣。目前的Sota语言模型,也就是。基于转换器的语言模型通常假设给定词的表示可以通过在凸壳中对其相关上下文的内插来捕获。然而,最近的研究表明,在高维空间中,几乎可以肯定的是,无论数据流形的潜在内在维度是什么,内插都不会发生。这种基于变换的语言模型生成的表示将收敛到一个密集的锥形超空间中,该超空间通常与许多不相邻的簇不连续。为了克服现有方法在大多数基于DL的自然语言处理模型中的局限性,本项目旨在利用勒贝格积分,通过自动识别这种不连续集合的边界来逼近给定输入单词特征的有限可测集合中的簇的后验分布,从而帮助生成更好的解释和量化不确定性。通过我们提出的基于勒贝格积分的近似,输入文本将由两个属性来表征:一个指示向量编码其在聚类(即,可测量集合)中的成员关系,以及另一个连续特征表示,以更好地捕捉其对下游任务的语义含义。这不仅允许在NLP中的输入文本的分布中更忠实地近似通常观察到的可数不连续,而且还使得能够学习更好地被人类理解的文本表示。
英文摘要
Deep learning (DL) based Natural Language Processing (NLP) technologies have attracted significant interest in recent years. The current SOTA language models, a.k.a. transformer-based language models, typically assume that the representation of a given word can be captured by the interpolation of its related context in a convex hull. However, it has recently been shown that in high-dimensional spaces, the interpolation almost surely never occurs regardless of the underlying intrinsic dimension of the data manifold. The representations generated by such transformer-based language models will converge into a dense cone-like hyperspace which is often discontinuous with many nonadjacent clusters. To overcome the limitation of current methods in most DL-based NLP models, this project aims to deploy Lebesgue integral, which can be defined as an ensemble of integrals among partitions (i.e., discontinuous feature clusters), to approximate the posterior distributions of clusters given input word features in finite measurable sets by automatically identifying the boundary of such discontinuous set, which in turn could help to generate better interpretations and quantify the uncertainty. By our proposed Lebesgue integral based approximation, the input text will be characterised by two properties: an indicator vector encoding its membership in clusters (i.e., measurable sets), and another continuous feature representation for better capturing its semantic meaning for downstream tasks. This not only allows for a more faithful approximation of commonly observed countably discontinuities in distributions of input text in NLP, but also enables learning text representations that are better understood by humans.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.48550/arxiv.2305.18926
发表时间: 2023-05
期刊:
影响因子: --
作者: [Xinyu Wang;Lin Gui;Yulan He]
通讯作者: Xinyu Wang;Lin Gui;Yulan He
Distilling ChatGPT for Explainable Automated Student Answer Assessment
提炼 ChatGPT 以进行可解释的自动化学生答案评估
DOI: 10.18653/v1/2023.findings-emnlp.399
发表时间: 2023
期刊:
影响因子: --
作者: [Li J]
通讯作者: Li J
DOI: 10.48550/arxiv.2306.03598
发表时间: 2023-06
期刊:
影响因子: --
作者: [Jiazheng Li;ZHAOYUE SUN;Bin Liang;Lin Gui;Yulan He]
通讯作者: Jiazheng Li;ZHAOYUE SUN;Bin Liang;Lin Gui;Yulan He
DOI: 10.48550/arxiv.2302.04985
发表时间: 2023-02
期刊: ArXiv
影响因子: --
作者: [Xingwei Tan;Gabriele Pergola;Yulan He]
通讯作者: Xingwei Tan;Gabriele Pergola;Yulan He
7
    国内基金
    海外基金
    用CLEAN和直接解调方法分析INTEGRAL数据
    • 批准号:
      10603004
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      35.0万元
    • 批准年份:
      2006
    • 负责人:
      周建锋
    • 依托单位: