A Mathematical View of Attention Models in Deep Learning

A Mathematical View of Attention Models in Deep Learning
复制标题

深度学习中注意力模型的数学观点

DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Yaochen Xie
Yaochen Xie
中科院分区:
--
文献类型:
--
作者:
Shuiwang Ji;Yaochen Xie

文献摘要

被引文献

相似文献

其中f(qj,ki)表征qJ和ki之间的关系(例如相似性),g(走气)通常是线性转换,因为g(vi)=wvvi∈R,wv∈RQ×p p,c = ∑i = ∑i = 1 f(qj,ki)是一个常规范围的范围。 θ(qj)φ(ki)),其中θ(走气)和φ(走气)通常是线性变换为θ(qj) = WQQJ和φ(Ki)= WKKI,如果我们将值向量作为输入,每个输出向量依赖于所有输入向量
where f(qj , ki) characterizes the relation (e.g., similarity) between qj and ki, g(⋅) is commonly a linear transformation as g(vi) = Wvvi ∈ R , where Wv ∈ Rq×p, and C = ∑i=1 f(qj , ki) is a normalization factor. A commonly used similarity function is the embedded Gaussian [5], defined as f(qj , ki) = exp (θ(qj)φ(ki)), where θ(⋅) and φ(⋅) are commonly linear transformations as θ(qj) = Wqqj and φ(ki) = Wkki. Note that if we treat the value vectors as inputs, each output vector oj is dependent on all input vectors. When the embedded Gaussian similarity and linear transformation are used, these computations can be expressed succinctly in matrix form as