papersTODAY 04:00 UTC
MAPS: Memory-Aware Predictive Scheduling for LLM Serving
Researchers propose MAPS, a scheduling framework designed to handle bursty large language model workloads on cloud infrastructure. The work targets memory-bound decode instances in prefill-decode disaggregated serving setups, where memory pressure limits throughput. It aims to improve scheduling decisions by predicting memory needs ahead of time.