benchmarks/tc_opt/test/parse_bench_results.py · 588538f5351d1fba33fa786d42d9fe795806987d · OpenDAS / vllm_cscc

"tests/kernels/attention/test_flash_attn.py" did not exist on "f256ebe4df6757d76f1f1642d7e110268a2f8190"

• feat(qwen3)：新增 vLLM 内置 RMS+RoPE 融合算子，并支持 LightOp 后端切换 · 588538f5

laibao authored Feb 06, 2026

  - 在 vLLM _C 扩展中新增 rms_rotary_embedding_fuse（注册 op + CUDA kernel），减少对 LightOp 的硬依赖
  - 新增环境变量 VLLM_FUSED_RMS_ROPE_BACKEND=auto|vllm|lightop，auto 优先走 vLLM，缺失时回退 LightOp
  - 更新 Qwen3 / Qwen3-MoE 的 fused 路径按后端选择执行
  - 补充 tc_opt benchmark 结果解析脚本 benchmarks/tc_opt/test/parse_bench_results.py

588538f5

parse_bench_results.py 6.8 KB

Replace parse_bench_results.py