[Model][Quantization] Fix / Add GGUF support for Qwen2 MoE models (#30307)

Signed-off-by: Tsukasa OI <floss_llm@irq.a4lg.com>

[Model][Quantization] Fix / Add GGUF support for Qwen2 MoE models (#30307)
Signed-off-by: Tsukasa OI <floss_llm@irq.a4lg.com>
73a484ca · Tsukasa OI · GitHub · b37bf51e · 73a484ca
Unverified Commit 73a484ca authored Dec 10, 2025 by Tsukasa OI Committed by GitHub Dec 09, 2025
Show whitespace changes
Inline Side-by-side

Showing with 8 additions and 0 deletions

vllm/model_executor/models/qwen2_moe.py vllm/model_executor/models/qwen2_moe.py +8 -0

No files found.
--- a/vllm/model_executor/models/qwen2_moe.py
+++ b/vllm/model_executor/models/qwen2_moe.py
@@ -367,6 +367,8 @@ class Qwen2MoeModel(nn.Module):
        self.embed_tokens = VocabParallelEmbedding(
            config.vocab_size,
            config.hidden_size,
+            quant_config=quant_config,
+            prefix=f"{prefix}.embed_tokens",
        )
        self.start_layer, self.end_layer, self.layers = make_layers(
            config.num_hidden_layers,
@@ -512,6 +514,12 @@ class Qwen2MoeModel(nn.Module):
                            continue
                        else:
                            name = remapped_kv_scale_name
+                    # GGUF: make sure that shared_expert_gate is a 2D tensor.
+                    if (
+                        "mlp.shared_expert_gate" in name
+                        and len(loaded_weight.shape) == 1
+                    ):
+                        loaded_weight = loaded_weight[None, :]
                    param = params_dict[name]
                    weight_loader = getattr(
                        param, "weight_loader", default_weight_loader