Merge branch 'v0.11.0-dev_tc_opt' into 'v0.11.0-dev'
perf(fused-moe): 预打包 Marlin W16A16 MoE 权重,降低 warmup 显存峰值 See merge request dcutoolkit/deeplearing/vllm!358
Showing
Please register or sign in to comment
perf(fused-moe): 预打包 Marlin W16A16 MoE 权重,降低 warmup 显存峰值 See merge request dcutoolkit/deeplearing/vllm!358