Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits

cs.AI updates on arXiv.org · 3h ago

arXiv:2609.38386v1 Announce Type: new Abstract: Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decoding. Fixed prefill chunks reduce this interference, but the best chunk size depends on the model, hardware, load, and latency objective. We introduce Decode-Latency Feedback…

Read original article on cs.AI updates on arXiv.org →