arXiv 2308.07661
Attention Is Not All You Need Anymore
By Zhe Chen
Published 2023-08-15
Wiki summary
Explore the paper's summary, context, and related research on Papiers.
In recent years, the popular Transformer architecture has achieved great success in many application areas, including natural language processing and computer vision. Many existing works aim to reduce the computational and memory complexity of the self-attention mechanism in the Transformer by trading off performance. However, performance is key for the continuing success of the Transformer. In this paper, a family…