课题基金 / 基金详情

Prostate cancer is a heterogeneous disease, displaying a multitude of genetic alterations, histological patterns and clinical outcomes. This heterogen

Prostate cancer is a heterogeneous disease, displaying a multitude of genetic alterations, histological patterns and clinical outcomes. This heterogen
前列腺癌是一种异质性疾病,表现出多种基因改变、组织学模式和临床结果。
批准号:
2432020
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该项目的主要目标是利用自然语言处理(NLP)的最新进展来开发端到端的临床支持系统,该系统可以利用电子健康记录(EHRs)中的纵向自由文本文档。电子病历通常包含历史记录,涉及患者与医疗保健服务之间的所有交互,包括免费文本文档,如转诊信和出院记录。一个值得注意的挑战是能够充分捕捉临床文本的纵向表征。常见的最先进的模型,如来自变形金刚的双向编码器表示(BERT)只能处理512个标记的序列(Devlin等人,2018),但单个患者一年的临床文本可以包含超过10,000个标记。另一个更普遍的问题与大型语言模型的透明度、可解释性和算法公平性有关。因此,本项目旨在制定方法和协议来增强这些方面。一种提出的表示顺序自由文本的方法建立在路径签名的基础上,这是一种以张量形式从数据中提取特征的非参数方法(Chevyrev和Kormilitzin, 2016)。广义地说,签名是关于时不变数据流的统计数据集合,具有普遍的非线性,因此它足以捕获原始数据的所有可能的非线性函数:允许一种独特的方法来表示复杂的顺序数据。将签名技术与策略相结合,以解决基于普通变压器的模型中注意机制的有限能力,例如空闲注意机制(Zaheer等人,2020)。这种混合方法应该允许有效的计算和患者临床文本历史的表示,可用于许多相关的下游任务。另一种方法将包含NLP研究中的新范式转变,称为提示学习。传统方法的许多下游任务都是采用BERT等模型对掩模语言建模(MLM)和下一句预测(NSP)进行预训练,然后对下游任务进行微调。而提示学习则重构预训练以嵌入下游任务,鼓励模型隐式学习期望的任务。在临床领域使用即时学习尚未被记录,因此提供了一个很大的机会。拟议的新方法将与临床医生协商制定和实施,并将处理实际的临床用例。具体来说,语言模型将在二级医疗UKCRIS数据库的大量自由文本笔记上进行训练,以帮助将患者分类到专家小组。其他部门将探索为临床试验识别病人和识别自我伤害的可行性。在EPSRC健康数据科学CDT提供的支持下,开发的方法和模型的翻译可行性将在心理健康范围之外进行测试。该项目属于EPSRC医疗保健技术研究领域。
英文摘要
The main aims of this project is to leverage recent advances in Natural Language Processing (NLP) to develop end-to-end clinical support systems which can utilise longitudinal free text documents within Electronic Health Records (EHRs). EHRs will often contain historic records, pertaining to all interactions between a patient and the healthcare service, including freetext documents, such as referral letters and discharge notes. A notable challenge is being able to adequately capturing longitudinal representations of clinical texts. Common state-of-the-art models such as the Bidirectional Encoder Representations fromTransformers (BERT) can only process sequences of 512 tokens (Devlin et al., 2018), but a years worth of clinical text for a single patient can consist of more than 10, 000 tokens. Another more general problem relates to the transparency, interpretability and algorithmic fairness of large language models. Therefore this project aims to develop methods and protocol to enhance these aspects.One proposed approach to representing sequential free-textbuilds upon the signature of a path, a non-parametric approach to extracting features from data in the form of tensors(Chevyrev and Kormilitzin, 2016). Loosely speaking, a signature is a collection of statistics about a stream of data that are time invariant, and has universal non-linearity, whereby it is sufficient to capture all possibly nonlinear functions of the original data: allowing aunique approach to representing complex sequential data. Combining signature techniques with strategies to address the limited ability of attention mechanisms in common transformer basedmodels, such as spare-attention mechanisms (Zaheer et al., 2020). This hybrid approach should allow efficient computation and representations of patients clinical text history, usable in a numberof relevant downstream tasks.Another approach will embrace a new paradigm shift in NLP research, named prompt-learning. Traditional approaches tomany downstream tasks involved taking a model such as BERT pre-trained on masked language modelling (MLM) and next sentence prediction (NSP) followed by a fine-tuning process on downstream tasks. Prompt-learning instead reconstructs the pretraining to embed the downstream task, encouraging the model to implicitly learn the desired task. The use of prompt-learning in a clinical domain has not been documented yet, thus provides a great opportunity.The proposed new methodologies will be developed and implemented in consultation with clinicians and will address real clinical use-cases. Specifically, the language models will be trained on a large collection of free-text notes from secondary care UKCRIS database to help triage patients to specialist teams. Other strands will explore the feasibility of identifying patients for clinical trials and identification of self-harm. The feasibility of translation of the developed methodology and models will be tested beyond the scope of mental health under the support provided by the EPSRC CDT in Health Data Science.This project falls within the EPSRC healthcare technologies research area.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
脊髓电刺激活化Na(V)1.1阳性GABA神经元持续缓解癌痛
  • 批准号:
    82371223
  • 项目类别:
    面上项目
  • 资助金额:
    49.00万元
  • 批准年份:
    2023
  • 负责人:
    闻大翔
  • 依托单位:
5'-tRF-GlyGCC通过SRSF1调控RNA可变剪切促三阴性乳腺癌作用机制及干预策略
  • 批准号:
    82372743
  • 项目类别:
    面上项目
  • 资助金额:
    49.00万元
  • 批准年份:
    2023
  • 负责人:
    陈卓佳
  • 依托单位:
丁酸梭菌代谢物(如丁酸、苯乳酸)通过MYC-TYMS信号轴影响结直肠癌化疗敏感性的效应及其机制研究
  • 批准号:
    82373139
  • 项目类别:
    面上项目
  • 资助金额:
    48.00万元
  • 批准年份:
    2023
  • 负责人:
    李孟鸿
  • 依托单位:
均相液相生物芯片检测系统的构建及其在癌症早期诊断上的应用
  • 批准号:
    82372089
  • 项目类别:
    面上项目
  • 资助金额:
    48.00万元
  • 批准年份:
    2023
  • 负责人:
    李万万
  • 依托单位: