Commits · 5f6d10c14c17122e6d711a4829ee0ca672e07f6f · OpenDAS / vllm_cscc

22 May, 2024 2 commits
- [CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722) · 5f6d10c1
  Michael Goin authored May 22, 2024
  
  5f6d10c1
- [Frontend] Dynamic RoPE scaling (#4638) · 9b9a10d6
  sasha0552 authored May 22, 2024
  
  9b9a10d6
21 May, 2024 7 commits
- [Bugfix][Kernel] Add head size check for attention backend selection (#4944) · 99eff67b
  Isotr0py authored May 22, 2024
  
  99eff67b
- [Bugfix] Fix flag name for `max_seq_len_to_capture` (#4935) · 14772eeb
  Kante Yin authored May 22, 2024
```
Signed-off-by: kerthcet <kerthcet@gmail.com>
```
  14772eeb
- [CI/Build] Codespell ignore `build/` directory (#4945) · 757b62c4
  Michael Goin authored May 21, 2024
  
  757b62c4
- [Docs] Add acknowledgment for sponsors (#4925) · e941f885
  Simon Mo authored May 21, 2024
  
  e941f885
- [Model] Add Phi-2 LoRA support (#4886) · f12c3b5b
  Isotr0py authored May 21, 2024
  
  f12c3b5b
- [Model] add rope_scaling support for qwen2 (#4930) · d130b573
  HUANG Fei authored May 21, 2024
  
  d130b573
- [Core] Fix scheduler considering "no LoRA" as "LoRA" (#4897) · 65ae8c2c
  Antoni Baum authored May 20, 2024
  
  65ae8c2c
20 May, 2024 8 commits
- [Doc]Add documentation to benchmarking script when running TGI (#4920) · c3af4472
  Kuntai Du authored May 20, 2024
  
  c3af4472
- [Core] Sharded State Loader download from HF (#4889) · 1937e298
  Aurick Qiao authored May 20, 2024
  
  1937e298
- [Bugfix] Fix dummy weight for fp8 (#4916) · f0eecee6
  Mor Zusman authored May 20, 2024
```
Allow dummy load format for fp8,
torch.uniform_ doesn't support FP8 at the moment
Co-authored-by: Mor Zusman <morz@ai21.com>
```
  f0eecee6
- [Build/CI] Enabling AMD Entrypoints Test (#4834) · 943e72ca
  Alexei-V-Ivanov-AMD authored May 20, 2024
```
Co-authored-by: Alexey Kondratiev <alexey.kondratiev@amd.com>
```
  943e72ca
- [Misc]: allow user to specify port in distributed setting (#4914) · 546a97ef
  Wenwei Zhang authored May 21, 2024
  
  546a97ef
- Remove marlin warning (#4918) · da5a0b53
  Alexander Matveev authored May 20, 2024
  
  da5a0b53
- [Model] LLaVA model refactor (#4910) · 6287537a
  Cyrus Leung authored May 20, 2024
  
  6287537a
- [Kernel] Add flash-attn back (#4907) · b57e6c59
  Woosuk Kwon authored May 19, 2024
  
  b57e6c59
19 May, 2024 2 commits
- [Kernel] Add marlin_24 unit tests (#4901) · 27ce8547
  Alexander Matveev authored May 19, 2024
  
  27ce8547
- [Bugfix][Model] Add base class for vision-language models (#4809) · f68470e8
  Cyrus Leung authored May 19, 2024
  
  f68470e8
18 May, 2024 2 commits

[Lora] Support long context lora (#4787) · 2e9a2227

SangBin Cho authored May 18, 2024

Currently we need to call rotary embedding kernel for each LoRA, which makes it hard to serve multiple long context length LoRA. Add batched rotary embedding kernel and pipe it through.

It replaces the rotary embedding layer to the one that is aware of multiple cos-sin-cache per scaling factors.

Follow up of https://github.com/vllm-project/vllm/pull/3095/files

2e9a2227

[ROCm][Hardware][AMD] Adding Navi21 to fallback to naive attention if Triton is not used (#4658) · c0724fc9
alexeykondrat authored May 18, 2024

c0724fc9

17 May, 2024 6 commits
- [Bugfix] Relax tiktoken to >= 0.6.0 (#4890) · 86b45ae0
  Michael Goin authored May 17, 2024
  
  86b45ae0
- [Doc] Update Ray Data distributed offline inference example (#4871) · c5711ef9
  Antoni Baum authored May 17, 2024
  
  c5711ef9
- Sync huggingface modifications of qwen Moe model (#4774) · 48d5985a
  eigenLiu authored May 18, 2024
  
  48d5985a
- [Bugfix] fix rope error when load models with different dtypes (#4835) · 33e0823d
  Jinzhen Lin authored May 17, 2024
  
  33e0823d
- [Build/CI] Extending the set of AMD tests with Regression, Basic Correctness,... · 26148120
  Alexei-V-Ivanov-AMD authored May 16, 2024
```
[Build/CI] Extending the set of AMD tests with Regression, Basic Correctness, Distributed, Engine, Llava Tests (#4797)
```
  26148120
- [Frontend] OpenAI API server: Do not add bos token by default when encoding (#4688) · 0150a106
  bofeng huang authored May 17, 2024
  
  0150a106
16 May, 2024 13 commits
- Support to serve vLLM on Kubernetes with LWS (#4829) · 8e7fb5d4
  Kante Yin authored May 17, 2024
```
Signed-off-by: kerthcet <kerthcet@gmail.com>
```
  8e7fb5d4
- [Bugfix] Fix FP8 KV cache support (#4869) · 9a31a817
  Woosuk Kwon authored May 16, 2024
  
  9a31a817
- [Kernel] Add w8a8 CUTLASS kernels (#4749) · 2060e936
  Tyler Michael Smith authored May 16, 2024
  
  2060e936
- [Kernel] Add punica dimension for Qwen1.5-32B LoRA (#4850) · 8435b207
  Silencio authored May 17, 2024
```
Co-authored-by: Silencio <silencio@adsl-99-6-187-6.dsl.irvnca.sbcglobal.net>
```
  8435b207
- [Misc] remove old comments (#4866) · 10fa9eea
  youkaichao authored May 16, 2024
  
  10fa9eea
- [Core][Distributed] remove graph mode function (#4818) · e0818808
  youkaichao authored May 16, 2024
  
  e0818808
- [ROCm][AMD][Bugfix] adding a missing triton autotune config (#4845) · b5853f99
  Hongxia Yang authored May 16, 2024
  
  b5853f99
- Add JSON output support for benchmark_latency and benchmark_throughput (#4848) · f09edd8a
  Simon Mo authored May 16, 2024
  
  f09edd8a
- Add GPTQ Marlin 2:4 sparse structured support (#4790) · 6979ade3
  Alexander Matveev authored May 16, 2024
```
Co-authored-by: Robert Shaw <rshaw@neuralmagic.com>
```
  6979ade3
- [Bugfix] Bypass authorization API token for preflight requests (#4862) · 9216b9cc
  Pierre Dulac authored May 16, 2024
  
  9216b9cc
- [Frontend] Separate OpenAI Batch Runner usage from API Server (#4851) · 5e0391c0
  Alex Wu authored May 16, 2024
  
  5e0391c0
- [docs] Fix typo in examples filename openi -> openai (#4864) · dbc0754d
  Alex Wu authored May 16, 2024
  
  dbc0754d
- [Kernel] add bfloat16 support for gptq marlin kernel (#4788) · 99caa491
  Jinzhen Lin authored May 16, 2024
  
  99caa491