Transformer models have demonstrated exceptional performance across a wide range of applications. Though forming the foundation of Transformer models, the dot-product attention does not scale well to long-context data since its time requirement grows quadratically with context length. In this work, we propose Radar, a training-free approach that accelerates inference by dynamically searching for the most important context tokens. For any pre-trained Transformer, Radar can reduce the decoding time complexity without training or heuristically evicting tokens. Moreover, we provide theoretical justification for our approach, demonstrating that Radar can reliably identify the most important tokens with high probability. We conduct extensive comparisons with the previous methods on a wide range of tasks. The results demonstrate that Radar achieves the state-of-the-art performance across different architectures with reduced time complexity, offering a practical solution for efficient long-context processing of Transformers.
Bibtex
@inproceedings{
hao2025radar,
title={Radar: Fast Long-Context Decoding for Any Transformer},
author={Yongchang Hao and Mengyao Zhai and Hossein Hajimirsadeghi and Sepidehsadat Hosseini and Frederick Tung},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=ZTpWOwMrzQ}
}
Related Research
-
Score-based Outlier Generation via Controlling the Radon-Nikodym Derivative
Score-based Outlier Generation via Controlling the Radon-Nikodym Derivative
A. Mukherjee, T. Milne, K. Y. C. Lui, S. Hazlewood, and J. Liu. IEEE Conference on Decision and Control
Publications
-
Investigating Action Embeddings for More Efficient Off-Policy Evaluation
Investigating Action Embeddings for More Efficient Off-Policy Evaluation
M. Wabartha, K. Wilson, R. David Evans, H. Sharifi, and T. Sylvain. RecSys CONSEQUENCES Workshop
Publications
-
FACTS: Fast, Accurate, and Privacy-Compliant TableSummarization via Offline Template Generation
FACTS: Fast, Accurate, and Privacy-Compliant TableSummarization via Offline Template Generation
Y. Yuan, M. Amin Shabani, and S. Liu. Workshop on Generative AI in Finance at NeurIPS
Publications