Enabling Deep Learning for Multilingual Sociopragmatics
Enabling Deep Learning for Multilingual Sociopragmatics
批准号:
RGPIN-2018-04267
负责人:
AbdulMageed, Muhammad
金额:
$2.04万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2019
资助国家:
加拿大
项目状态:
已结题
起止时间:
2019-01-01 至 2020-12-31
中文摘要
自然语言处理(NLP)是一个令人兴奋的领域,专注于教计算机理解和生成人类语言。近年来,深度学习作为一种受人脑信息处理启发的机器学习方法,在许多具有大量标注数据的自然语言处理任务(如机器翻译、语音识别)上都打破了记录。由于这些进步及其带来的普及技术,自然语言的深度学习目前是一个具有很高社会经济影响的战略领域。这使得创建在社会语用学层面上理解人类语言的模型的时机已经成熟(即,一个命题的意义取决于它所在的社会语境)。然而,许多挑战依然存在。两个突出的、相互关联的例子是(A)与标记数据相关的高成本,以及(B)现有标记数据中的偏差。缺乏标记数据阻碍了构建强大的深度学习模型的进展,因为这些模型只在给定大量标记数据的情况下进行缩放。有偏见的标签数据导致创建的技术比其他技术更好地服务于特定的主导群体,这可能会产生严重的社会和经济影响。*我的研究计划旨在开发方法,通过针对这两个核心问题,在社会语用学水平上加速自然语言的深度学习,重点是将自然语言处理技术带到几种语言和语言变体的更广泛的人口结构中。*该提案有三个关键目标:*1.跨语言代用标签:这涉及为社会语用任务(例如,用户意图建模、同理心检测)开发自动标签数据的方法,重点是英语和所有阿拉伯变体(即代表所有22个阿拉伯国家的变体)。*2.深度生成性半监督学习:我将开发利用深度生成性模型的方法,深度生成性模型是一类深度学习方法,可以生成可用作标签数据的合理语言。这将有助于解决上述两个以数据为重点的问题(即a和b)。*3.具有受控社交语用学的社交机器:我的目标是开发能够根据熟人属性(例如,情感智能语言生成、特定性别和个性的对话代理)进行动态定制的社交语料学智能对话模型。*这项研究将在各个领域有广泛的应用,包括决策、健康和福祉、教育、娱乐和娱乐。由于它是一个专业子领域,处于一些已经受到供应限制的领域的交界处,自然语言的深度学习目前面临着严重的人才短缺。我的项目提供的HQP培训将有助于满足这些不断增长的需求。
英文摘要
Natural language processing (NLP) is the exciting field focused at teaching computers to understand and generate human language. Recently, deep learning, a class of machine learning methods inspired by information processing in the human brain, has broken records on many NLP tasks for which large amounts of labeled data are available (e.g., machine translation, speech recognition). Due to these advances and the pervasive technologies it enables, deep learning of natural language is currently a strategic area of high socioeconomic impact. This makes it ripe time for creating models that understand human language at the level of sociopragmatics (i.e., the meaning of a proposition depends on the social context in which it is uttered). Many challenges, however, remain. Two prominent, inter-related, examples are (a) the high costs associated to labeling data, and (b) the bias in existing labeled data. Absence of labeled data hinders progress on building powerful deep learning models since these models scale exclusively given large amounts of labeled data. Biased labeled data result in creating technologies that serve particular dominant groups better than others, which can have serious social and economic repercussions.******My research program aims at developing methods to accelerate deep learning of natural language at the level of sociopragmatics by targeting these two core problems, with a focus on bringing NLP technologies to wider demographics across several languages and language varieties. ******The proposal has three key objectives: ******1. Cross-Lingual Surrogate Labeling: This involves developing methods for automatically labeling data for sociopragmatic tasks (e.g., user intention modeling, empathy detection), with a focus on English and all Arabic varieties (i.e., varieties representing all the 22 Arab countries). ******2. Deep Generative Semi-Supervised Learning: I will develop methods that exploit deep generative models, a class of deep learning methods that can generate sensible language that can be leveraged as labeled data. This will help solve the two data-focused problems above (i.e., a and b). ******3. Toward Social Machines With Controlled Sociopragmatics: My goal is to develop sociopragmatically intelligent conversational models capable of dynamic customization in response to conversant attributes (e.g., emotionally intelligent language generation, gender- and personality-specific conversational agents).******The research will have a wide range of applications in various fields, including decision making, health and well-being, education, recreation, and entertainment. Since it is a specialized subfield at the junction of a number of already supply-constrained fields, deep learning of natural language currently suffers from acute shortage of talent. HQP training provided by my program will contribute to fulfilling these ever-rising needs.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Enabling Deep Learning for Multilingual Sociopragmatics
-
批准号:RGPIN-2018-04267
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2022
-
负责人:AbdulMageed, Muhammad
-
依托单位:
Enabling Deep Learning for Multilingual Sociopragmatics
-
批准号:RGPIN-2018-04267
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2021
-
负责人:AbdulMageed, Muhammad
-
依托单位:
Enabling Deep Learning for Multilingual Sociopragmatics
-
批准号:RGPIN-2018-04267
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2020
-
负责人:AbdulMageed, Muhammad
-
依托单位:
Enabling Deep Learning for Multilingual Sociopragmatics
-
批准号:RGPIN-2018-04267
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2018
-
负责人:AbdulMageed, Muhammad
-
依托单位:
Enabling Deep Learning for Multilingual Sociopragmatics
-
批准号:DGECR-2018-00369
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2018
-
负责人:AbdulMageed, Muhammad
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Deep Seek引导下预防肝硬化腹水患者发生腹腔感染的约翰霍普金斯循证实践模型下中医护理策略的构建研究
-
批准号:2026JJ81909
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:胡曦
-
依托单位:
基于Deep Unrolling的高分辨近红外二区荧光分子断层成像方法研究
-
批准号:12271434
-
项目类别:面上项目
-
资助金额:46万元
-
批准年份:2022
-
负责人:贺小伟
-
依托单位:
基于深度森林(Deep Forest)模型的表面增强拉曼光谱分析方法研究
-
批准号:2020A151501709
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2020
-
负责人:谢怡
-
依托单位:
面向Deep Web的数据整合关键技术研究
-
批准号:61872168
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2018
-
负责人:董永权
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于语义计算的海量Deep Web知识探索机制研究
-
批准号:61272411
-
项目类别:面上项目
-
资助金额:80.0万元
-
批准年份:2012
-
负责人:赵峰
-
依托单位:
Deep Web数据集成查询结果抽取与整合关键技术研究
-
批准号:61100167
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2011
-
负责人:董永权
-
依托单位:
面向Deep Web的大规模知识库自动构建方法研究
-
批准号:61170020
-
项目类别:面上项目
-
资助金额:57.0万元
-
批准年份:2011
-
负责人:崔志明
-
依托单位:
Deep Web敏感聚合信息保护方法研究
-
批准号:61003054
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2010
-
负责人:赵朋朋
-
依托单位: