
FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving
arXiv:2610.06917v1 Announce Type: new Abstract: Prefill-decode disaggregation is becoming a common architecture for LLM serving because it separates two phases with distinct execution patterns and SLO objectives. Existing systems typically combine a fixed prefill/decode worker ratio with request routing across workers. However, real-world workloads exhibit both short bursts and sustained shifts…
Read original article on cs.AI updates on arXiv.org →