Natural language processing in healthcare data
Natural language processing in healthcare data
批准号:
RGPIN-2019-04701
负责人:
Rudzicz, Frank
金额:
$1.68万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2019
资助国家:
加拿大
项目状态:
已结题
起止时间:
2019-01-01 至 2020-12-31
中文摘要
词嵌入(即,“词向量”或“分布式表示”)是词的密集数字表示,其用作各种统计机器学习方法的输入。通常情况下,通过优化上下文统计,这些嵌入诱导潜在的维度,编码形态,语法,甚至语义方面。因此,结果可以捕获传统方法无法提供的数据中概念之间的有意义的关系。Vector Institute正在与临床评价科学研究所(ICES)合作,共同使用EMRALD语料库,该语料库由来自各种初级保健来源的文本组成(例如,咨询记录、转诊、风险因素、既往病史),这些信息来自安大略的数百名医生。EMRALD在词汇量和总体规模上都比Google的新闻语料库大一个数量级,Google的新闻语料库是用于训练嵌入的事实上的语料库之一。目前,极大的词汇量似乎产生了两个主要后果:i)技术术语及其许多变体的优势,以及B)拼写错误。这些结果导致非常稀疏的上下文矩阵。**我们在这个研究计划中有三个主要目标:*1)用本体信息丰富词嵌入。我们的团队已经开发出一种使用多任务学习方法和来自众包统计数据的规范词汇数据来“丰富”嵌入的方法。例如,用情感规范丰富嵌入过程不仅提高了情感分析的准确性,而且还提高了域外任务的准确性,例如,机器翻译在这里,我们打算采用类似的方法,但结构化的本体信息从医学文本和资源。* *审计分类器所做的决策,并确保其各自模型中个人信息的隐私变得越来越重要。例如,最近表明,可以使用另一组最小链接数据重新识别匿名数据集中的患者。为了提高模型的可解释性,我们将应用LIME、基于文本的“锚定”和差分隐私等方法。我们将探讨生成对抗网络是否也可以合成具有类似属性的分布。**3)进行纵向分类。最初的目标是使用EMRALD中的结构化数据对给定临床记录的诊断代码进行监督分类。考虑到数据的纵向性质,这将包括具有注意力的递归神经网络和卷积神经网络。我们将通过删除一些结构化数据或向标签添加噪声来类似地探索半监督学习。长期目标是将这些方法联合收割机结合起来,以预测各种长期趋势和人类轨迹。
英文摘要
Word embeddings (i.e., 'word vectors' or 'distributed representations') are dense numeric representations of words, which serve as input to various statistical machine learning methods. Typically, by optimizing contextual statistics, these embeddings induce latent dimensions that encode aspects of morphology, syntax, and even semantics. The results therefore can capture meaningful relationships among concepts in the data not afforded by traditional methods.******The Vector Institute is partnering with the Institute for Clinical Evaluative Sciences (ICES) around the collaborative use of the EMRALD corpus, which consists of text from a variety of primary care sources (e.g., consult notes, referrals, risk factors, past medical history) sourced from hundreds of doctors in Ontario. EMRALD is an order of magnitude larger, in vocabulary and overall size, than Google's news corpus which is one of the de facto corpora used for training embeddings. Currently, the extremely large vocabulary size appears to produce two main consequences: i) a preponderance of technical terms and their many variants, and b) spelling mistakes. These consequences lead to very sparse contextual matrices.******We have three primary goals in this program of research:******1) To enrich word embeddings with ontological information. Our team has developed a method of 'enriching' embeddings using a multi-task learning approach and normative lexical data, from crowd-sourced statistics. For example, enriching the embedding process with norms of sentiment increases the accuracy not only of sentiment analysis, but out--of--domain tasks as well, e.g., machine translation. Here, we intend to apply a similar approach but with structured ontological information from medical texts and resources. ******2) To produce explainable and private models. It is increasingly important to audit decisions made by classifiers, and to ensure the privacy of personal information in their respective models. For instance, it was recently shown that it is possible to re-identify patients in an anonymized data set using another set of minimally linked data. In order to increase the explainability of our models, we will apply methods such as LIME , text--based 'anchoring', and differential privacy. We will explore whether generative adversarial networks can also synthesize distributions with similar properties. ******3) To perform longitudinal classification. An initial goal will be to use the structured data in EMRALD to perform supervised classification of diagnostic codes given clinical notes. Given the longitudinal nature of the data, this will include recurrent neural networks and convolutional neural networks with attention. We will similarly explore semi-supervised learning either by removing some structured data or adding noise to the labels. The long-term aim is to combine these approaches in order to predict various long-term trends and human trajectories.**
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Machine learning in surgical safety
-
批准号:RGPIN-2020-05910
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$0.72万
-
财政年份:2022
-
负责人:Rudzicz, Frank
-
依托单位:
Machine learning in surgical safety
-
批准号:RGPIN-2020-05910
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.03万
-
财政年份:2022
-
负责人:Rudzicz, Frank
-
依托单位:
Machine learning in surgical safety
-
批准号:RGPIN-2020-05910
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2021
-
负责人:Rudzicz, Frank
-
依托单位:
Machine learning in surgical safety
-
批准号:RGPIN-2020-05910
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2020
-
负责人:Rudzicz, Frank
-
依托单位:
Analyzing Child Language Experiences Around the World (ACLEW)
-
批准号:501769-2016
-
项目类别:Discovery Frontiers - Digging into Data
-
资助金额:$1.17万
-
财政年份:2018
-
负责人:Rudzicz, Frank
-
依托单位:
A control-theoretic model of speech production and recognition for use within prosthetic communication devices
-
批准号:435874-2013
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2018
-
负责人:Rudzicz, Frank
-
依托单位:
Automatic remote screening of speech features associated with Alzheimer's disease
-
批准号:508463-2017
-
项目类别:Collaborative Health Research Projects
-
资助金额:$11.02万
-
财政年份:2018
-
负责人:Rudzicz, Frank
-
依托单位:
Automatic remote screening of speech features associated with Alzheimer's disease
-
批准号:508463-2017
-
项目类别:Collaborative Health Research Projects
-
资助金额:$5.01万
-
财政年份:2017
-
负责人:Rudzicz, Frank
-
依托单位:
A control-theoretic model of speech production and recognition for use within prosthetic communication devices
-
批准号:435874-2013
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2017
-
负责人:Rudzicz, Frank
-
依托单位:
Exploiting natural neural control to minimize speech errors among language learners
-
批准号:522798-2017
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2017
-
负责人:Rudzicz, Frank
-
依托单位:
Analyzing Child Language Experiences Around the World (ACLEW)
-
批准号:501769-2016
-
项目类别:Discovery Frontiers - Digging into Data
-
资助金额:$5.1万
-
财政年份:2017
-
负责人:Rudzicz, Frank
-
依托单位:
Analyzing Child Language Experiences Around the World (ACLEW)
-
批准号:501769-2016
-
项目类别:Discovery Frontiers - Digging into Data
-
资助金额:$1.02万
-
财政年份:2016
-
负责人:Rudzicz, Frank
-
依托单位:
A control-theoretic model of speech production and recognition for use within prosthetic communication devices
-
批准号:435874-2013
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2016
-
负责人:Rudzicz, Frank
-
依托单位:
System for recording and analyzing telephone conversations between humans and artificial intelligence
-
批准号:RTI-2017-00811
-
项目类别:Research Tools and Instruments
-
资助金额:$1.17万
-
财政年份:2016
-
负责人:Rudzicz, Frank
-
依托单位:
A control-theoretic model of speech production and recognition for use within prosthetic communication devices
-
批准号:435874-2013
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2015
-
负责人:Rudzicz, Frank
-
依托单位:
Machine learning for discovering conversation strategies and recovering from breakdowns
-
批准号:476952-2014
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2014
-
负责人:Rudzicz, Frank
-
依托单位:
A control-theoretic model of speech production and recognition for use within prosthetic communication devices
-
批准号:435874-2013
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2014
-
负责人:Rudzicz, Frank
-
依托单位:
Automated assessment of customer service through voice recognition
-
批准号:474704-2014
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2014
-
负责人:Rudzicz, Frank
-
依托单位:
A control-theoretic model of speech production and recognition for use within prosthetic communication devices
-
批准号:435874-2013
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.46万
-
财政年份:2013
-
负责人:Rudzicz, Frank
-
依托单位:
Speaker identification for use in modern hearing aids
-
批准号:452764-2013
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2013
-
负责人:Rudzicz, Frank
-
依托单位:
国内基金
海外基金
登录
查看更多内容
儿童音乐能力发展对语言与社会认知能力及脑发育的影响
-
批准号:31971003
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:南云
-
依托单位:
面向英汉双向跨语言图像检索的文本分析关键技术研究
-
批准号:61170095
-
项目类别:面上项目
-
资助金额:57.0万元
-
批准年份:2011
-
负责人:张玥杰
-
依托单位:
儿童植入耳蜗后听觉行为与言语发展进程的关联性研究
-
批准号:81170916
-
项目类别:面上项目
-
资助金额:65.0万元
-
批准年份:2011
-
负责人:刘莎
-
依托单位:
基于儿童心理分析的图解式汉语口语自动解析方法研究
-
批准号:60175012
-
项目类别:面上项目
-
资助金额:18.0万元
-
批准年份:2001
-
负责人:宗成庆
-
依托单位: