CRII: RI: Using Linguistic Variation to Understand Deep Neural Models of Language
CRII: RI: Using Linguistic Variation to Understand Deep Neural Models of Language
批准号:
2139005
负责人:
Kyle Mahowald
金额:
$17.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-07-01 至 2024-06-30
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Many successful modern computational language systems rely on deep neural networks. Whereas older techniques relied on structured linguistic representations, the inner workings of neural models can be opaque even to the engineers and scientists who create them. Therefore, a current major challenge in Natural Language Processing, as in other areas of Artificial Intelligence, is to develop methods that allow us to understand the internal representations of opaque neural models. Natural Language Processing is uniquely poised to contribute to this endeavor for two reasons. First, the field of linguistics has long sought to develop tools for characterizing the kinds of representations necessary for processing human language, and so there is a rich body of prior work to draw on. Second, the breadth and variation of world languages give us a natural way of studying models under different, but equally valid, parameterizations. Just as studying how humans can process diverse languages gives insight into human language processing and human cognition, understanding how multilingual computational systems process different languages can give insights into computational models.A class of deep neural models, known as transformers, has been particularly successful at natural language tasks. Some of these models are massively multilingual, trained on large numbers of languages at once. Interestingly, these multilingual models seem to acquire both language-specific and language-general knowledge. Taking advantage of linguistic techniques and variation among world languages, this Computer Research Initiation Research (CRII) project undertakes a series of computational experiments that involve training small classifiers on the pre-trained embedding space of massive multilingual models (e.g., Multilingual BERT and XLM-Roberta) and using the classifier output to characterize how these models represent crucial grammatical aspects of language (e.g., grammatical subject) across languages with different morphosyntactic systems. Moreover, in order to develop more robust ways of studying grammatical roles in these models, the project uses computational techniques to build and publicly release more richly annotated multilingual corpora. The experimental results and public corpora contribute both to our understanding of computational language models and diversify the set of languages that can be studied using these techniques.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.18653/v1/2021.emnlp-main.471
发表时间:
2021-09
期刊:
影响因子:
--
作者:
[Alex Jones;W. Wang;Kyle Mahowald]
通讯作者:
Alex Jones;W. Wang;Kyle Mahowald
DOI:
10.18653/v1/2021.eacl-main.215
发表时间:
2021-01
期刊:
ArXiv
影响因子:
--
作者:
[Isabel Papadimitriou;Ethan A. Chi;Richard Futrell;Kyle Mahowald]
通讯作者:
Isabel Papadimitriou;Ethan A. Chi;Richard Futrell;Kyle Mahowald
What do tokens know about their characters and how do they know it?
令牌对它们的角色了解多少?它们是如何知道的?
DOI:
10.18653/v1/2022.naacl-main.179
发表时间:
2022
期刊:
Proceedings of NAACL
影响因子:
--
作者:
[Kaushal, Ayush, Mahowald, Kyle]
通讯作者:
Mahowald, Kyle
When classifying grammatical role, BERT doesn’t care about word order... except when it matters
在对语法角色进行分类时,BERT 并不关心词序……除非它很重要
DOI:
10.18653/v1/2022.acl-short.71
发表时间:
2022
期刊:
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers
影响因子:
--
作者:
[Papadimitriou, Isabel, Futrell, Richard, Mahowald, Kyle]
通讯作者:
Mahowald, Kyle
longhorns at DADC 2022: How many linguists does it take to fool a Question Answering model? A systematic approach to adversarial attacks.
DADC 2022 上的长角牛:需要多少语言学家才能愚弄问答模型?
DOI:
10.18653/v1/2022.dadc-1.5
发表时间:
2022
期刊:
Proceedings of DADC Workshop
影响因子:
--
作者:
[Kovatchev, Venelin, Chatterjee, Trina, Govindarajan, Venkata S, Chen, Jifan, Choi, Eunsol, Chronis, Gabriella, Das, Anubrata, Erk, Katrin, Lease, Matthew, Li, Junyi Jessy]
通讯作者:
Li, Junyi Jessy
CAREER: Investigating linguistic and cognitive abstractions for solving word problems in minds and machines
-
批准号:2339729
-
项目类别:Continuing Grant
-
资助金额:$136.9万
-
财政年份:2024
-
负责人:Kyle Mahowald
-
依托单位:
CRII: RI: Using Linguistic Variation to Understand Deep Neural Models of Language
-
批准号:2104995
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2021
-
负责人:Kyle Mahowald
-
依托单位:
国内基金
海外基金
登录
查看更多内容
破骨细胞源性FcγRI介导类风湿性关节炎炎症后疼痛的作用机制
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:阳林
-
依托单位:
四神丸调控生物钟基因Bmal1/Fc εRI介导肥大细胞节律性活化治疗IBS-D“晨起痛”的作用机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:何心凌
-
依托单位:
NSUN6介导的m5C修饰调控心肌细胞凋亡和铁死亡参与MI/RI的机制研究
-
批准号:2026JJ80739
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:袁乐宏
-
依托单位:
中药牛耳枫中抗MI/RI新颖虎皮楠生物碱的发现与作用机制研究
-
批准号:2026JJ60255
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:张济辉
-
依托单位:
醒脑静多靶点调控PI3K/Akt通路抑制CI/RI氧化应激—基于网络药理学及体内、外实验研究
-
批准号:2025JJ90117
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:李秋云
-
依托单位:
IgA-FcαRI介导的Syk/NLRP3/caspase-1通路在线状IgA大疱性皮病
中的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:荆可
-
依托单位:
基于双修饰ANG-RNH1系统阻抑RI复合物生成机制建立口腔黏膜等效物血管化稳态
-
批准号:82401112
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2024
-
负责人:刘旭倩
-
依托单位:
跨膜蛋白LRP5胞外域调控膜受体TβRI促钛表面BMSCs归巢、分化的研究
-
批准号:82301120
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2023
-
负责人:於科
-
依托单位:
基于“免疫-神经”网络探讨眼针活化CI/RI大鼠MC靶向H3R调节“免疫监视”的抗炎机制
-
批准号:82374375
-
项目类别:面上项目
-
资助金额:51万元
-
批准年份:2023
-
负责人:马贤德
-
依托单位:
Dectin-2通过促进FcεRI聚集和肥大细胞活化加剧哮喘发作的机制研究
-
批准号:82300022
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2023
-
负责人:屈玉兰
-
依托单位: