
RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
arXiv:2609.20971v1 Announce Type: new Abstract: Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse block selection can reduce this cost, but a block centroid may hide a highly relevant token among many irrelevant ones. We call this failure mode mean dilution and propose…
Read original article on cs.AI updates on arXiv.org →