Heaps do lie: debugging a memory leak in vLLM.
Mistral AI investigated a vLLM memory leak in a disaggregated Prefill/Decode deployment of Mistral Medium 3.1. System memory grew linearly at 400 MB per minute, causing out-of-memory failures within hours. It occurred only on the decode side with graph compilation and KVCache transfer via NIXL. Heaptrack showed stable heap memory; the issue was outside the heap.