课题基金 / 基金详情

A Lebesgue Integral based Approximation for Language Modelling

A Lebesgue Integral based Approximation for Language Modelling
基于勒贝格积分的语言建模近似
批准号:
EP/X019063/1
负责人:
Lin Gui
金额:
$25.77万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
近年来,基于深度学习(DL)的自然语言处理(NLP)技术引起了人们的极大兴趣。当前的SOTA语言模型,也就是基于变换器的语言模型通常假定给定单词的表示可以通过在凸船体中对其相关上下文的内插来捕获。然而,最近已经表明,在高维空间中,插值几乎肯定不会发生,无论数据流形的内在维度。由这种基于转换器的语言模型生成的表示将收敛到一个密集的锥状超空间中,该超空间通常是不连续的,具有许多不相邻的簇。为了克服大多数基于DL的NLP模型中当前方法的局限性,该项目旨在部署Lebesgue积分,它可以被定义为分区之间的积分集合(即,不连续特征聚类),通过自动识别这种不连续集合的边界来近似有限可测量集合中给定输入词特征的聚类的后验分布,这反过来可以帮助生成更好的解释并量化不确定性。通过我们提出的基于Lebesgue积分的近似,输入文本将由两个属性来表征:一个指示符向量,其编码其在聚类中的成员资格(即,可测量集合),以及另一个连续特征表示,用于更好地捕获其语义含义以用于下游任务。这不仅可以更忠实地近似NLP中输入文本分布中常见的可数不连续性,而且还可以学习人类更好理解的文本表示。
英文摘要
Deep learning (DL) based Natural Language Processing (NLP) technologies have attracted significant interest in recent years. The current SOTA language models, a.k.a. transformer-based language models, typically assume that the representation of a given word can be captured by the interpolation of its related context in a convex hull. However, it has recently been shown that in high-dimensional spaces, the interpolation almost surely never occurs regardless of the underlying intrinsic dimension of the data manifold. The representations generated by such transformer-based language models will converge into a dense cone-like hyperspace which is often discontinuous with many nonadjacent clusters. To overcome the limitation of current methods in most DL-based NLP models, this project aims to deploy Lebesgue integral, which can be defined as an ensemble of integrals among partitions (i.e., discontinuous feature clusters), to approximate the posterior distributions of clusters given input word features in finite measurable sets by automatically identifying the boundary of such discontinuous set, which in turn could help to generate better interpretations and quantify the uncertainty. By our proposed Lebesgue integral based approximation, the input text will be characterised by two properties: an indicator vector encoding its membership in clusters (i.e., measurable sets), and another continuous feature representation for better capturing its semantic meaning for downstream tasks. This not only allows for a more faithful approximation of commonly observed countably discontinuities in distributions of input text in NLP, but also enables learning text representations that are better understood by humans.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.48550/arxiv.2305.18926
发表时间: 2023-05
期刊:
影响因子: --
作者: [Xinyu Wang;Lin Gui;Yulan He]
通讯作者: Xinyu Wang;Lin Gui;Yulan He
Distilling ChatGPT for Explainable Automated Student Answer Assessment
提炼 ChatGPT 以进行可解释的自动化学生答案评估
DOI: 10.18653/v1/2023.findings-emnlp.399
发表时间: 2023
期刊:
影响因子: --
作者: [Li J]
通讯作者: Li J
DOI: 10.48550/arxiv.2306.03598
发表时间: 2023-06
期刊:
影响因子: --
作者: [Jiazheng Li;ZHAOYUE SUN;Bin Liang;Lin Gui;Yulan He]
通讯作者: Jiazheng Li;ZHAOYUE SUN;Bin Liang;Lin Gui;Yulan He
DOI: 10.48550/arxiv.2302.04985
发表时间: 2023-02
期刊: ArXiv
影响因子: --
作者: [Xingwei Tan;Gabriele Pergola;Yulan He]
通讯作者: Xingwei Tan;Gabriele Pergola;Yulan He
7
    国内基金
    海外基金
    用CLEAN和直接解调方法分析INTEGRAL数据
    • 批准号:
      10603004
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      35.0万元
    • 批准年份:
      2006
    • 负责人:
      周建锋
    • 依托单位: