llm: Support KV cache quantization with gpt-oss
With the new version of GGML in #12245, KV cache quantization no longer causes a fallback to CPU.
Showing
Please register or sign in to comment
With the new version of GGML in #12245, KV cache quantization no longer causes a fallback to CPU.