Commits · db9e5708a98b7209cf4465a0391139cf8fca7674 · OpenDAS / vllm_cscc

29 Jul, 2024 1 commit
- [Kernel] Tuned FP8 Kernels for Ada Lovelace (#6677) · 766435e6
  Varun Sundar Rabindranath authored Jul 29, 2024
```
Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>
```
  766435e6
27 Jul, 2024 3 commits
- [Kernel] Increase precision of GPTQ/AWQ Marlin kernel (#6795) · 75acdaa4
  Alexander Matveev authored Jul 27, 2024
  
  75acdaa4
- [Model] H2O Danube3-4b (#6451) · 14dbd5a7
  Joe authored Jul 26, 2024
  
  14dbd5a7
- [Bug Fix] Illegal memory access, FP8 Llama 3.1 405b (#6852) · 55712941
  Lucas Wilkinson authored Jul 26, 2024
  
  55712941
26 Jul, 2024 2 commits
- [Hardware] [Intel] Enable Multiprocessing and tensor parallel in CPU backend... · 3bbb4936
  Li, Jiang authored Jul 27, 2024
```
[Hardware] [Intel] Enable Multiprocessing and tensor parallel in CPU backend and update documentation  (#6125)
```
  3bbb4936
- [Bugfix][Kernel] Promote another index to int64_t (#6838) · 50704f52
  Tyler Michael Smith authored Jul 26, 2024
  
  50704f52
24 Jul, 2024 1 commit
- Add fp8 support to `reshape_and_cache_flash` (#6667) · 0e63494c
  Antoni Baum authored Jul 24, 2024
  
  0e63494c
22 Jul, 2024 1 commit
- [Bugfix][Kernel] Use int64_t for indices in fp8 quant kernels (#6649) · fea59c77
  Tyler Michael Smith authored Jul 22, 2024
  
  fea59c77
21 Jul, 2024 1 commit
- [Kernel][Core] Add AWQ support to the Marlin kernel (#6612) · 396d92d5
  Alexander Matveev authored Jul 21, 2024
  
  396d92d5
20 Jul, 2024 1 commit
- [ Kernel ] FP8 Dynamic Per Token Quant - Add scale_ub (#6593) · 2e265642
  Varun Sundar Rabindranath authored Jul 19, 2024
```
Co-authored-by: Varun Sundar Rabindranth <varun@neuralmagic.com>
```
  2e265642
18 Jul, 2024 1 commit
- [ Kernel ] FP8 Dynamic-Per-Token Quant Kernel (#6511) · b5241e41
  Varun Sundar Rabindranath authored Jul 17, 2024
```
Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>
```
  b5241e41
17 Jul, 2024 1 commit
- [Core] draft_model_runner: Implement prepare_inputs on GPU for advance_step (#6338) · e76466dd
  Alexander Matveev authored Jul 17, 2024
  
  e76466dd
16 Jul, 2024 1 commit
- [Kernel][Attention] Separate `Attention.kv_scale` into `k_scale` and `v_scale` (#6081) · 978aed53
  Michael Goin authored Jul 16, 2024
  
  978aed53
14 Jul, 2024 1 commit
- [Kernel] Turn off CUTLASS scaled_mm for Ada Lovelace (#6384) · 9dad5cc8
  Tyler Michael Smith authored Jul 14, 2024
  
  9dad5cc8
03 Jul, 2024 1 commit
- [Kernel] Expand FP8 support to Ampere GPUs using FP8 Marlin (#5975) · 47f0954a
  Michael Goin authored Jul 03, 2024
  
  47f0954a
29 Jun, 2024 1 commit
- [Kernel] Add punica dimensions for Granite 3b and 8b (#5930) · ba499444
  Joe Runde authored Jun 28, 2024
```
Signed-off-by: Joe Runde <joe@joerun.de>
```
  ba499444
28 Jun, 2024 2 commits
- Unmark more files as executable (#5962) · 5d2a1a9c
  Tyler Michael Smith authored Jun 28, 2024
  
  5d2a1a9c
- [Bugfix] Fix compute datatype for cutlass 3.x epilogues (#5931) · 6a2d659d
  Tyler Michael Smith authored Jun 28, 2024
  
  6a2d659d
26 Jun, 2024 2 commits
- Support CPU inference with VSX PowerPC ISA (#5652) · 38a1674a
  Chip Kerchner authored Jun 26, 2024
  
  38a1674a
- [Kernel] Adding bias epilogue support for `cutlass_scaled_mm` (#5560) · 5bfd1bbc
  Luka Govedič authored Jun 26, 2024
```
Co-authored-by: Chih-Chieh-Yang <7364402+cyang49@users.noreply.github.com>
Co-authored-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
```
  5bfd1bbc
23 Jun, 2024 1 commit
- [BugFix] [Kernel] Add Cutlass2x fallback kernels (#5744) · 6c916ac8
  Varun Sundar Rabindranath authored Jun 24, 2024
```
Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>
```
  6c916ac8
21 Jun, 2024 2 commits
- [Kernel][CPU] Add Quick `gelu` to CPU (#5717) · bd620b01
  Roger Wang authored Jun 20, 2024
  
  bd620b01
- [Kernel] Add punica dimension for Qwen2 LoRA (#5441) · 1f567421
  Jinzhen Lin authored Jun 21, 2024
  
  1f567421
20 Jun, 2024 4 commits
- [Bugfix] Fix the CUDA version check for FP8 support in the CUTLASS kernels (#5715) · 3f3b6b21
  Tyler Michael Smith authored Jun 20, 2024
  
  3f3b6b21
- [Kernel] Update Cutlass int8 kernel configs for SM80 (#5275) · a7dcc620
  Varun Sundar Rabindranath authored Jun 20, 2024
```
Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>
```
  a7dcc620
- [Model] Port over CLIPVisionModel for VLMs (#5591) · ad137cd1
  Roger Wang authored Jun 20, 2024
  
  ad137cd1
- [Kernel] Update Cutlass int8 kernel configs for SM90 (#5514) · 111af1fa
  Varun Sundar Rabindranath authored Jun 20, 2024
```
Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>
```
  111af1fa
18 Jun, 2024 3 commits
- [Bugfix] Fix CUDA version check for mma warning suppression (#5642) · b23ce920
  Tyler Michael Smith authored Jun 18, 2024
  
  b23ce920
- [Model] LoRA support added for command-r (#5178) · 07feecde
  sergey-tinkoff authored Jun 18, 2024
  
  07feecde
- [Kernel] Add punica dimensions for Granite 13b (#5559) · 5002175e
  Joe Runde authored Jun 17, 2024
```
Signed-off-by: Joe Runde <Joseph.Runde@ibm.com>
```
  5002175e
14 Jun, 2024 2 commits
- [Kernel] Suppress mma.sp warning on CUDA 12.5 and later (#5401) · 348616ac
  Tyler Michael Smith authored Jun 14, 2024
  
  348616ac
- [Kernel] Fix CUTLASS 3.x custom broadcast load epilogue (#5516) · 703475f6
  Tyler Michael Smith authored Jun 14, 2024
  
  703475f6
13 Jun, 2024 2 commits

[Hardware][Intel] Support CPU inference with AVX2 ISA (#5452) · cd9c0d65
Jie Fu (傅杰) authored Jun 14, 2024

cd9c0d65

[Kernel] Factor out epilogues from cutlass kernels (#5391) · 85657b56

Tyler Michael Smith authored Jun 13, 2024


Co-authored-by: Michael Goin <michael@neuralmagic.com>
Co-authored-by: youkaichao <youkaichao@gmail.com>
Co-authored-by: zifeitong <zifei.tong@parasail.io>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-neuralmagic@users.noreply.github.com>

85657b56

12 Jun, 2024 1 commit

[Kernel] Vectorized FP8 quantize kernel (#5396) · 5985e342

Cody Yu authored Jun 12, 2024

Inspired by #5146, this PR improves FP8 quantize kernel by vectorizing data transfer to better utilize memory bandwidth. Microbenchmark shows that this improved kernel can achieve 1.0x-1.5x speedup (especially when hidden size is large).

In details, we applied 3 optimizations:

- Use inverted scale so that most divisions are changed to multiplications.
- Unroll the loop by 4 times to improve ILP.
- Use vectorized 4 to transfer data between HBM and SRAM.

5985e342

09 Jun, 2024 1 commit
- [Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047) · 5467ac31
  bnellnm authored Jun 09, 2024
  
  5467ac31
07 Jun, 2024 2 commits
- [Misc] Remove unused cuda_utils.h in CPU backend (#5345) · 6840a716
  Jie Fu (傅杰) authored Jun 08, 2024
  
  6840a716
- [Kernel] Dynamic Per-Token Activation Quantization (#5037) · ca3ea51b
  Dipika Sikka authored Jun 07, 2024
```
Co-authored-by: Varun Sundar Rabindranath <varunsundar08@gmail.com>
Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>
```
  ca3ea51b
05 Jun, 2024 1 commit
- [Kernel] Add GPU architecture guards to the CUTLASS w8a8 kernels to reduce binary size (#5157) · ccd4f129
  Tyler Michael Smith authored Jun 05, 2024
```
Co-authored-by: Cody Yu <hao.yu.cody@gmail.com>
```
  ccd4f129
03 Jun, 2024 1 commit
- [CI/BUILD] enable intel queue for longer CPU tests (#4113) · cafb8e06
  Yuan authored Jun 04, 2024
  
  cafb8e06