Early Exiting BERT for Efficient Document Ranking
Early Exiting BERT for Efficient Document Ranking
复制标题
DOI:
10.18653/v1/2020.sustainlp-1.11
复制
发表时间:
2020-11
期刊:
影响因子:
--
通讯作者:
Ji Xin;Rodrigo Nogueira;Yaoliang Yu;Jimmy J. Lin
中科院分区:
文献类型:
--
作者:
Ji Xin;Rodrigo Nogueira;Yaoliang Yu;Jimmy J. Lin
Pre-trained language models such as BERT have shown their effectiveness in various tasks. Despite their power, they are known to be computationally intensive, which hinders real-world applications. In this paper, we introduce early exiting BERT for document ranking. With a slight modification, BERT becomes a model with multiple output paths, and each inference sample can exit early from these paths. In this way, computation can be effectively allocated among samples, and overall system latency is significantly reduced while the original quality is maintained. Our experiments on two document ranking datasets demonstrate up to 2.5x inference speedup with minimal quality degradation. The source code of our implementation can be found at https://github.com/castorini/earlyexiting-monobert.