Performance Le cache KV (Key-Value) expliqué Comment le cache KV accélère l'inférence, comment calculer sa taille, et comment le compresser : GQA, MLA, quantification q8_0/fp8, SWA, PagedAttention. 1 juin 2026 · Antoine Michéa