Attention-Aware Routing: Coupling Routing and Attention in MoEs

cs.AI updates on arXiv.org · 2d ago
Research Papers

arXiv:2609.20974v1 Announce Type: new Abstract: In Mixture-of-Experts language models, the router typically selects and weights experts based on the token's hidden state, utilizing limited contextual information. We propose Attention-Aware Routing (AAR), which augments the router with temporal and spectral features extracted from a sliding window of attention weights that represent a summary of…

Read original article on cs.AI updates on arXiv.org →