A Mathematical View of Attention Models in Deep Learning
A Mathematical View of Attention Models in Deep Learning
复制标题
深度学习中注意力模型的数学观点
DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Yaochen Xie
中科院分区:
文献类型:
--
作者:
Shuiwang Ji;Yaochen Xie
where f(qj , ki) characterizes the relation (e.g., similarity) between qj and ki, g(⋅) is commonly a linear transformation as g(vi) = Wvvi ∈ R , where Wv ∈ Rq×p, and C = ∑i=1 f(qj , ki) is a normalization factor. A commonly used similarity function is the embedded Gaussian [5], defined as f(qj , ki) = exp (θ(qj)φ(ki)), where θ(⋅) and φ(⋅) are commonly linear transformations as θ(qj) = Wqqj and φ(ki) = Wkki. Note that if we treat the value vectors as inputs, each output vector oj is dependent on all input vectors. When the embedded Gaussian similarity and linear transformation are used, these computations can be expressed succinctly in matrix form as