Latency-SLO-Aware Memory Offloading for Large Language Model Inference

Document Type

Conference Proceeding

Publication Date

7-5-2026

Abstract

Offloading large language models (LLMs) states to host memory during inference promises to reduce operational costs by supporting larger models, longer prompts, and larger batch sizes. However, the design of existing memory offloading mechanisms does not take latency service-level objectives (SLOs) into consideration. As a result, they either lead to frequent SLO violations or underutilize host memory, thereby incurring economic loss and thus defeating the purpose of memory offloading. This paper presents NovaServe, a latency-SLO-aware memory offloading system for LLM serving. A key challenge in designing NovaServe is to reconcile the tension between meeting SLOs and maximizing host memory usage. To address this, NovaServe introduces a dual-dimensional offloading strategy that jointly controls the weight offloading ratio and attention computation placement to capture the precise tradeoff between SLOs and host memory usage. With that, NovaServe proposes a two-stage approach to automatically pick the strategy. The first stage is offline which generates the optimal offloading ratios for certain SLOs, while the second stage adjusts the strategy at the granularity of inference iteration based on runtime hardware status. Our evaluation shows that NovaServe consistently meets SLOs and improves the serving throughput over existing mechanisms by 2.3 × through efficiently utilizing host memory.

Publication Title

Proceedings of the International Conference on Supercomputing

ISBN

[9798400725227]

Share

COinS