Mistral AI·· 2026-01-22
Heaps do lie: debugging a memory leak in vLLM.
Heaps do lie: debugging a memory leak in vLLM.
AI summary
Mistral AI investigated a vLLM memory leak in a disaggregated Prefill/Decode deployment of Mistral Medium 3.1. System memory grew linearly at 400 MB per minute, causing out-of-memory failures within hours. It occurred only on the decode side with graph compilation and KVCache transfer via NIXL. Heaptrack showed stable heap memory; the issue was outside the heap.
Selection record
Threshold 60Official, first-handFirst 34Second 34
Not admittedSum of both 68 < twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:vLLM推理内存泄漏调试,AI工程实践
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Mistral AI · mistral.ai