Commits · 96b6f475dda40a0c7d557f73c36fe09c07be2e9c · kecinstone / 2024pra-vllm

01 Feb, 2024 1 commit
- Remove hardcoded `device="cuda" ` to support more devices (#2503) · 96b6f475
  Kunshang Ji authored Feb 02, 2024
```
Co-authored-by: Jiang Li <jiang1.li@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
```
  96b6f475
31 Jan, 2024 1 commit
- Add unit test for Mixtral MoE layer (#2677) · d0d93b92
  Philipp Moritz authored Jan 31, 2024
  
  d0d93b92
30 Jan, 2024 1 commit
- DeepseekMoE support with Fused MoE kernel (#2453) · 5d60def0
  wangding zeng authored Jan 30, 2024
```
Co-authored-by: roy <jasonailu87@gmail.com>
```
  5d60def0
29 Jan, 2024 1 commit

Support FP8-E5M2 KV Cache (#2279) · 9090bf02

zhaoyang-star authored Jan 29, 2024


Co-authored-by: zhaoyang <zhao.yang16@zte.com.cn>
Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>

9090bf02

27 Jan, 2024 1 commit
- AWQ: Up to 2.66x higher throughput (#2566) · beb89f68
  Casper authored Jan 27, 2024
  
  beb89f68
23 Jan, 2024 1 commit

[Experimental] Add multi-LoRA support (#1804) · 9b945daa

Antoni Baum authored Jan 24, 2024


Co-authored-by: Chen Shen <scv119@gmail.com>
Co-authored-by: Shreyas Krishnaswamy <shrekris@anyscale.com>
Co-authored-by: Avnish Narayan <avnish@anyscale.com>

9b945daa

18 Jan, 2024 1 commit

[Experimental] Prefix Caching Support (#1669) · d10f8e1d

shiyi.c_98 authored Jan 17, 2024


Co-authored-by: DouHappy <2278958187@qq.com>
Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>

d10f8e1d

15 Jan, 2024 1 commit
- fix weigit loading for GQA with TP (#2379) · f780504d
  Chenhui Zhang authored Jan 16, 2024
  
  f780504d
12 Jan, 2024 1 commit

Aligning `top_p` and `top_k` Sampling (#1885) · 218dc2cc

陈序 authored Jan 13, 2024

* Align top_p and top_k with huggingface

* remove _get_prompt_and_output_tokens

* rename _apply_top_p_top_k

* compare top_p top_k with hf

* fix test errors

218dc2cc

09 Jan, 2024 1 commit
- [Speculative decoding 1/9] Optimized rejection sampler (#2336) · 79d64c49
  Cade Daniel authored Jan 09, 2024
  
  79d64c49
08 Jan, 2024 1 commit
- [Minor] Remove unused code in attention (#2384) · 28c3f121
  Woosuk Kwon authored Jan 08, 2024
  
  28c3f121
03 Jan, 2024 2 commits
- Use NCCL instead of ray for control-plane communication to remove serialization overhead (#2221) · fd4ea8ef
  Zhuohan Li authored Jan 04, 2024
  
  fd4ea8ef
- [Minor] Fix typo and remove unused code (#2305) · 91405610
  Roy authored Jan 03, 2024
  
  91405610
20 Dec, 2023 1 commit
- Remove Sampler copy stream (#2209) · bd29cf3d
  Antoni Baum authored Dec 20, 2023
  
  bd29cf3d
17 Dec, 2023 2 commits
- Make sampler less blocking (#1889) · a7347d9a
  Antoni Baum authored Dec 17, 2023
  
  a7347d9a
- Optimize model execution with CUDA graph (#1926) · 37ca5581
  Woosuk Kwon authored Dec 16, 2023
```
Co-authored-by: Chen Shen <scv119@gmail.com>
Co-authored-by: Antoni Baum <antoni.baum@protonmail.com>
```
  37ca5581
15 Dec, 2023 1 commit
- Add GPTQ support (#916) · 0fbfc4b8
  CHU Tianxiang authored Dec 15, 2023
  
  0fbfc4b8
12 Dec, 2023 1 commit
- Support MPT with GQA (#1938) · 6428f1d0
  Megha Agarwal authored Dec 12, 2023
```
Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>
```
  6428f1d0
10 Dec, 2023 1 commit
- Replace head_mapping params with num_kv_heads to attention kernel. (#1997) · dacaf5a4
  wbn authored Dec 11, 2023
```
Co-authored-by: wangguoya <wangguoya@baidu.com>
Co-authored-by: Yang Zhao <zhaoyangstar@foxmail.com>
```
  dacaf5a4
08 Dec, 2023 1 commit

Merge EmbeddedLLM/vllm-rocm into vLLM main (#1836) · 6ccc0bff

TJian authored Dec 08, 2023


Co-authored-by: Philipp Moritz <pcmoritz@gmail.com>
Co-authored-by: Amir Balwel <amoooori04@gmail.com>
Co-authored-by: root <kuanfu.liu@akirakan.com>
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: kuanfu <kuanfu.liu@embeddedllm.com>
Co-authored-by: miloice <17350011+kliuae@users.noreply.github.com>

6ccc0bff

03 Dec, 2023 1 commit
- Add PyTorch-native implementation of custom layers (#1898) · 9b294976
  Woosuk Kwon authored Dec 02, 2023
  
  9b294976
30 Nov, 2023 3 commits
- Fix rope cache key error (#1867) · d27f4bae
  Roy authored Dec 01, 2023
  
  d27f4bae
- Avoid multiple instantiations of the RoPE class (#1828) · 63b2206a
  Jee Li authored Nov 30, 2023
  
  63b2206a
- Refactor Worker & InputMetadata (#1843) · 27feead2
  Woosuk Kwon authored Nov 29, 2023
  
  27feead2
29 Nov, 2023 1 commit
- Refactor Attention (#1840) · a9e45742
  Woosuk Kwon authored Nov 29, 2023
  
  a9e45742
28 Nov, 2023 1 commit
- [FIX] Fix class naming (#1803) · 708e6c18
  Zhuohan Li authored Nov 28, 2023
  
  708e6c18
24 Nov, 2023 1 commit
- [Build] Avoid building too many extensions (#1624) · e0c6f556
  Yanming W authored Nov 23, 2023
  
  e0c6f556
22 Nov, 2023 1 commit
- Fix repetition penalty aligned with huggingface (#1577) · de23687d
  ljss authored Nov 23, 2023
  
  de23687d
21 Nov, 2023 3 commits
- [FIX] Fix the case when `input_is_parallel=False` for `ScaledActivation` (#1737) · 7d761fe3
  Zhuohan Li authored Nov 20, 2023
  
  7d761fe3
- [BugFix] Fix TP support for AWQ (#1731) · cf35d8f3
  Woosuk Kwon authored Nov 20, 2023
  
  cf35d8f3
- Rewrite torch.repeat_interleave to remove cpu synchronization (#1599) · 819b18e7
  ljss authored Nov 21, 2023
  
  819b18e7
20 Nov, 2023 1 commit
- Migrate linter from `pylint` to `ruff` (#1665) · 5ffc0d13
  Simon Mo authored Nov 20, 2023
  
  5ffc0d13
19 Nov, 2023 2 commits
- [Optimization] Implement fused add rmsnorm (#1667) · e1054247
  ljss authored Nov 19, 2023
  
  e1054247
- Add AWQ support for all models (#1714) · 8d17774f
  Woosuk Kwon authored Nov 18, 2023
  
  8d17774f
18 Nov, 2023 1 commit
- Support Min P Sampler (#1642) · e87557b0
  Roy authored Nov 18, 2023
  
  e87557b0
16 Nov, 2023 1 commit

TP/quantization/weight loading refactor part 2 - Refactor quantized linear... · 7076fa1c

Zhuohan Li authored Nov 15, 2023

TP/quantization/weight loading refactor part 2 - Refactor quantized linear logic and extend quantization support to all models (#1622)

Refactor the tensor parallelism, quantization, and weight-loading codes.

Summary of the new features enabled by this PR:
- **All models** are able to be quantized with AWQ and SqueezeLLM, and [soon GPTQ](https://github.com/vllm-project/vllm/pull/1580).
- Model loading code became much simpler.
- Support model parallelism for all MQA/GQA models when the number of key/value heads is smaller than the tensor parallel size.

7076fa1c

13 Nov, 2023 1 commit
- [Minor] Move RoPE selection logic to `get_rope` (#1633) · 054072be
  Woosuk Kwon authored Nov 12, 2023
  
  054072be
03 Nov, 2023 2 commits
- Support YaRN models (#1264) · 9f669a9a
  Antoni Baum authored Nov 03, 2023
```
Signed-off-by: Antoni Baum <antoni.baum@protonmail.com>
Co-authored-by: Viktor Ferenczi <viktor@ferenczi.eu>
Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>
```
  9f669a9a
- Added logits processor API to sampling params (#1469) · 555bdcc5
  Noam Gat authored Nov 03, 2023
  
  555bdcc5
01 Nov, 2023 1 commit
- Force paged attention v2 for long contexts (#1510) · 9738b84a
  Antoni Baum authored Nov 01, 2023
  
  9738b84a