RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

cs.AI updates on arXiv.org · 2d ago
Research Papers

arXiv:2609.20971v1 Announce Type: new Abstract: Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse block selection can reduce this cost, but a block centroid may hide a highly relevant token among many irrelevant ones. We call this failure mode mean dilution and propose…

Read original article on cs.AI updates on arXiv.org →