FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving

cs.AI updates on arXiv.org · 2h ago

arXiv:2610.06917v1 Announce Type: new Abstract: Prefill-decode disaggregation is becoming a common architecture for LLM serving because it separates two phases with distinct execution patterns and SLO objectives. Existing systems typically combine a fixed prefill/decode worker ratio with request routing across workers. However, real-world workloads exhibit both short bursts and sustained shifts…

Read original article on cs.AI updates on arXiv.org →