Skip to content
Mistral AI·· 2026-01-22

Heaps do lie: debugging a memory leak in vLLM.

Heaps do lie: debugging a memory leak in vLLM.

AI summary

Mistral AI investigated a vLLM memory leak in a disaggregated Prefill/Decode deployment of Mistral Medium 3.1. System memory grew linearly at 400 MB per minute, causing out-of-memory failures within hours. It occurred only on the decode side with graph compilation and KVCache transfer via NIXL. Heaptrack showed stable heap memory; the issue was outside the heap.

Selection record

Not admittedSum of both 68 < twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:vLLM推理内存泄漏调试,AI工程实践

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Mistral AI · mistral.ai