RI: Medium: Learning Disentangled Representations for Text to Aid Interpretability and Transfer
RI: Medium: Learning Disentangled Representations for Text to Aid Interpretability and Transfer
批准号:
1901117
负责人:
Byron Wallace
金额:
$100.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-07-01 至 2024-06-30
中文摘要
用于自然语言处理的机器学习方法支持我们日常使用的许多技术,例如垃圾邮件过滤器和翻译软件。这些技术背后的模型已经变得越来越复杂,带来了更好的性能,但也增加了复杂性。特别是,基于“神经网络”的方法已经重新成为语言处理的机器学习模型的主导类别。这些方法通常比非神经方法表现得更好,但也有关键的缺点。首先,训练这些模型需要人力和时间来生成足够大的人工标注文本形式的训练数据集。其次,在一个数据集上训练的模型是否会推广到另一个数据集通常并不明显。最后,很难看出为什么这些模型会做出它们所做的具体预测,这在很大程度上是因为预测是基于对文本的习得表示做出的,而文本本身并不具备透明度。该项目提出了技术创新,以解决这些相互关联的问题,使用“解缠”。其想法是设计模型,使用于做出预测的学习表示法具有已知的意义。这种方法有可能实现模型的重用(提高效率和降低人力成本),并有助于解释,以便人们可以更好地了解为什么模型做出给定的预测。为了实现上述提高模型的可解释性和可转移性的目标,本工作将开发和评估学习特定维度被注入显式语义的表示的新模型。这与当前的方法不同,当前的方法不分青红皂白地将所有属性编码到一个(纠缠的)表示中。为了实现解缠,该项目将探索深度生成模型和稀疏、门控神经编码器。这些将使用归纳偏差和光监督策略,引导模型朝着解开纠缠的表示方向发展。例如,如果学习的嵌入空间中的距离不反映人类对实例相对于感兴趣的特定方面的相对相似性的判断,则模型将受到惩罚。在其他情况下,“薄弱的”监督(例如规则)可为解除纠缠提供适当的指导。最后,“探查”任务构成了需要探索的第三种监督策略:这将涉及到使用辅助任务来提供“监督”,以引导个人从各个方面嵌入输入。该项目将为自然语言处理中的代表性问题开发和评估这些模型,特别是:分类、序列标记和摘要。模型将被评估为预测性能(包括它们对新领域的概括性和这样做的效率),以及学习到的表示被解开并捕捉到预期方面的程度。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Machine learning methods for natural language processing power many technologies that we use on a day-to-day basis, such as spam filters and translation software. The models underlying these techniques have become increasingly sophisticated, yielding improved performance but also increasing complexity. In particular, "neural network" based approaches have re-emerged as the dominant class of machine learning models for language processing. These approaches often perform better than their non-neural counterparts, but also have key downsides. First, training these models requires human effort and time to generate a sufficiently large set of training data in the form of manually annotated text. Second, it is often not obvious whether a model trained on one dataset will generalize to another. Finally, it is hard to discern why such models make the specific predictions that they do, largely because predictions are made on the basis of learned representations of texts which do not naturally afford transparency. This project proposes technical innovations to address these interrelated issues using "disentanglement". The idea is to design models such that the learned representations used to make predictions have known meaning. This approach has the potential to enable re-use of models (increasing efficiency and reducing human costs), and aid interpretability, so that one can have a better idea of why a model made a given prediction.To realize the above goals of improved interpretability and transferability of models, this work will develop and evaluate new models that learn representations in which certain dimensions are imbued with explicit semantics. This is a departure from current approaches, which indiscriminately code all attributes into a single (entangled) representation. To achieve disentanglement, this project will explore deep generative models and sparse, gated neural encoders. These will use inductive biases and light supervision strategies that guide models toward disentangled representations. For example, models will be penalized if distances in learned embedding spaces do not reflect human judgments concerning the relative similarities of instances with respect to specific aspects of interest. In other cases, "weak" supervision (e.g., rules) may provide adequate guidance for disentanglement. Finally, "probing" tasks constitute a third supervision strategy to be explored: This will involve the use of auxiliary tasks to provide "supervision" that guides individual aspect-wise embeddings of input. The project will develop and evaluate such models for representative problems in natural language processing, specifically: classification, sequence tagging, and summarization. Models will be evaluated both for predictive performance (including their generalizability to new domains and the efficiency with which they do so), and the degree to which learned representations are disentangled and capture the intended aspects.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1162/coli_a_00397
发表时间:
2021-03-01
期刊:
COMPUTATIONAL LINGUISTICS
影响因子:
9.3
作者:
[Agarwal, Oshin, Yang, Yinfei, Nenkova, Ani]
通讯作者:
Nenkova, Ani
Rate-Regularization and Generalization in Variational Autoencoders
变分自编码器中的速率正则化和泛化
DOI:
--
发表时间:
2021
期刊:
Proceedings of The 24th International Conference on Artificial Intelligence and Statistics
影响因子:
--
作者:
[Bozkurt, A, Esmaeili, B., Tristan, J.-B., Brooks, D., Dy, J., van de Meent, J.-W.]
通讯作者:
van de Meent, J.-W.
DOI:
10.48550/arxiv.2210.06565
发表时间:
2022-10
期刊:
Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing
影响因子:
--
作者:
[Denis Jered McInerney;Geoffrey S. Young;Jan-Willem van de Meent;Byron Wallace]
通讯作者:
Denis Jered McInerney;Geoffrey S. Young;Jan-Willem van de Meent;Byron Wallace
Biomedical Interpretable Entity Representations
生物医学可解释的实体表示
DOI:
--
发表时间:
2021
期刊:
Proceedings of the Association for Computational Linguistics (ACL
影响因子:
--
作者:
[Garcia-Olano, Diego, Onoe, Yasumasa, Baldini, Ioana, Ghosh, Joydeep, Wallace, Byron C., Varshney, Kush]
通讯作者:
Varshney, Kush
DOI:
10.18653/v1/2022.findings-acl.153
发表时间:
2021-07
期刊:
ArXiv
影响因子:
--
作者:
[Pouya Pezeshkpour;Sarthak Jain;Sameer Singh;Byron C. Wallace]
通讯作者:
Pouya Pezeshkpour;Sarthak Jain;Sameer Singh;Byron C. Wallace
共 9 条
Collaborative Research: RI: Medium: Expert-in-the-Loop Neural Summarization for Consequential Domains
-
批准号:2211954
-
项目类别:Standard Grant
-
资助金额:$60.0万
-
财政年份:2022
-
负责人:Byron Wallace
-
依托单位:
CAREER: Structured Scientific Evidence Extraction: Models and Corpora
-
批准号:1750978
-
项目类别:Continuing Grant
-
资助金额:$54.99万
-
财政年份:2018
-
负责人:Byron Wallace
-
依托单位:
Collaborative research: ABI Development: Making Advanced Statistical Tools Accessible for Quantitative Research Synthesis and Discovery in Ecology and Evolutionary Biology
-
批准号:1520781
-
项目类别:Standard Grant
-
资助金额:$25.59万
-
财政年份:2014
-
负责人:Byron Wallace
-
依托单位:
Collaborative research: ABI Development: Making Advanced Statistical Tools Accessible for Quantitative Research Synthesis and Discovery in Ecology and Evolutionary Biology
-
批准号:1262442
-
项目类别:Standard Grant
-
资助金额:$50.27万
-
财政年份:2013
-
负责人:Byron Wallace
-
依托单位:
海外基金