课题基金 / 基金详情

A Lebesgue Integral based Approximation for Language Modelling

A Lebesgue Integral based Approximation for Language Modelling
基于勒贝格积分的语言建模近似
批准号:
EP/X019063/1
负责人:
Lin Gui
金额:
$25.77万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
近年来,基于深度学习(DL)的自然语言处理(NLP)技术引起了人们的极大兴趣。当前的SOTA语言模型,也称为基于转换器的语言模型,通常假设给定单词的表示可以通过在凸包中插入相关上下文来捕获。然而,最近的研究表明,在高维空间中,无论数据流形的内在维数是多少,插值几乎肯定不会发生。由这种基于变换的语言模型生成的表示将收敛成密集的锥状超空间,该空间通常是不连续的,有许多不相邻的簇。为了克服大多数基于dl的NLP模型中现有方法的局限性,本项目旨在部署Lebesgue积分,该积分可以定义为分区(即不连续特征聚类)之间的积分集合,通过自动识别这种不连续集的边界来近似有限可测量集中给定输入词特征的聚类的后测分布,从而有助于生成更好的解释和量化不确定性。通过我们提出的基于勒贝格积分的近似,输入文本将由两个属性来表征:一个指示向量编码其在聚类(即可测量集)中的隶属关系,另一个连续特征表示用于更好地捕获其下游任务的语义。这不仅可以更忠实地近似NLP中输入文本分布中通常观察到的可数不连续,而且还可以学习人类更好理解的文本表示。
英文摘要
Deep learning (DL) based Natural Language Processing (NLP) technologies have attracted significant interest in recent years. The current SOTA language models, a.k.a. transformer-based language models, typically assume that the representation of a given word can be captured by the interpolation of its related context in a convex hull. However, it has recently been shown that in high-dimensional spaces, the interpolation almost surely never occurs regardless of the underlying intrinsic dimension of the data manifold. The representations generated by such transformer-based language models will converge into a dense cone-like hyperspace which is often discontinuous with many nonadjacent clusters. To overcome the limitation of current methods in most DL-based NLP models, this project aims to deploy Lebesgue integral, which can be defined as an ensemble of integrals among partitions (i.e., discontinuous feature clusters), to approximate the posterior distributions of clusters given input word features in finite measurable sets by automatically identifying the boundary of such discontinuous set, which in turn could help to generate better interpretations and quantify the uncertainty. By our proposed Lebesgue integral based approximation, the input text will be characterised by two properties: an indicator vector encoding its membership in clusters (i.e., measurable sets), and another continuous feature representation for better capturing its semantic meaning for downstream tasks. This not only allows for a more faithful approximation of commonly observed countably discontinuities in distributions of input text in NLP, but also enables learning text representations that are better understood by humans.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.48550/arxiv.2305.18926
发表时间: 2023-05
期刊:
影响因子: --
作者: [Xinyu Wang;Lin Gui;Yulan He]
通讯作者: Xinyu Wang;Lin Gui;Yulan He
Distilling ChatGPT for Explainable Automated Student Answer Assessment
提炼 ChatGPT 以进行可解释的自动化学生答案评估
DOI: 10.18653/v1/2023.findings-emnlp.399
发表时间: 2023
期刊:
影响因子: --
作者: [Li J]
通讯作者: Li J
DOI: 10.48550/arxiv.2306.03598
发表时间: 2023-06
期刊:
影响因子: --
作者: [Jiazheng Li;ZHAOYUE SUN;Bin Liang;Lin Gui;Yulan He]
通讯作者: Jiazheng Li;ZHAOYUE SUN;Bin Liang;Lin Gui;Yulan He
DOI: 10.48550/arxiv.2302.04985
发表时间: 2023-02
期刊: ArXiv
影响因子: --
作者: [Xingwei Tan;Gabriele Pergola;Yulan He]
通讯作者: Xingwei Tan;Gabriele Pergola;Yulan He
7
    国内基金
    海外基金
    用CLEAN和直接解调方法分析INTEGRAL数据
    • 批准号:
      10603004
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      35.0万元
    • 批准年份:
      2006
    • 负责人:
      周建锋
    • 依托单位: