Commits · de23687d168ebeaa8872c27f05b8292bab0fac71 · kecinstone / 2024pra-vllm

22 Nov, 2023 1 commit
- Fix repetition penalty aligned with huggingface (#1577) · de23687d
  ljss authored Nov 23, 2023
  
  de23687d
21 Nov, 2023 3 commits
- [FIX] Fix the case when `input_is_parallel=False` for `ScaledActivation` (#1737) · 7d761fe3
  Zhuohan Li authored Nov 20, 2023
  
  7d761fe3
- [BugFix] Fix TP support for AWQ (#1731) · cf35d8f3
  Woosuk Kwon authored Nov 20, 2023
  
  cf35d8f3
- Rewrite torch.repeat_interleave to remove cpu synchronization (#1599) · 819b18e7
  ljss authored Nov 21, 2023
  
  819b18e7
20 Nov, 2023 1 commit
- Migrate linter from `pylint` to `ruff` (#1665) · 5ffc0d13
  Simon Mo authored Nov 20, 2023
  
  5ffc0d13
19 Nov, 2023 2 commits
- [Optimization] Implement fused add rmsnorm (#1667) · e1054247
  ljss authored Nov 19, 2023
  
  e1054247
- Add AWQ support for all models (#1714) · 8d17774f
  Woosuk Kwon authored Nov 18, 2023
  
  8d17774f
18 Nov, 2023 1 commit
- Support Min P Sampler (#1642) · e87557b0
  Roy authored Nov 18, 2023
  
  e87557b0
16 Nov, 2023 1 commit

TP/quantization/weight loading refactor part 2 - Refactor quantized linear... · 7076fa1c

Zhuohan Li authored Nov 15, 2023

TP/quantization/weight loading refactor part 2 - Refactor quantized linear logic and extend quantization support to all models (#1622)

Refactor the tensor parallelism, quantization, and weight-loading codes.

Summary of the new features enabled by this PR:
- **All models** are able to be quantized with AWQ and SqueezeLLM, and [soon GPTQ](https://github.com/vllm-project/vllm/pull/1580).
- Model loading code became much simpler.
- Support model parallelism for all MQA/GQA models when the number of key/value heads is smaller than the tensor parallel size.

7076fa1c

13 Nov, 2023 1 commit
- [Minor] Move RoPE selection logic to `get_rope` (#1633) · 054072be
  Woosuk Kwon authored Nov 12, 2023
  
  054072be
03 Nov, 2023 2 commits
- Support YaRN models (#1264) · 9f669a9a
  Antoni Baum authored Nov 03, 2023
```
Signed-off-by: Antoni Baum <antoni.baum@protonmail.com>
Co-authored-by: Viktor Ferenczi <viktor@ferenczi.eu>
Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>
```
  9f669a9a
- Added logits processor API to sampling params (#1469) · 555bdcc5
  Noam Gat authored Nov 03, 2023
  
  555bdcc5
01 Nov, 2023 1 commit
- Force paged attention v2 for long contexts (#1510) · 9738b84a
  Antoni Baum authored Nov 01, 2023
  
  9738b84a
30 Oct, 2023 1 commit
- Delay GPU->CPU sync in sampling (#1337) · 15f56323
  Antoni Baum authored Oct 30, 2023
  
  15f56323
29 Oct, 2023 1 commit
- Support repetition_penalty (#1424) · 69be658b
  ljss authored Oct 30, 2023
  
  69be658b
22 Oct, 2023 1 commit

Support SqueezeLLM (#1326) · 1f24755b

chooper1 authored Oct 22, 2023


Co-authored-by: squeeze-ai-lab <squeezeailab.bair@gmail.com>
Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>

1f24755b

17 Oct, 2023 1 commit
- Change scheduler & input tensor shape (#1381) · c1376e0f
  Woosuk Kwon authored Oct 16, 2023
  
  c1376e0f
16 Oct, 2023 2 commits
- Implement prompt logprobs & Batched topk for computing logprobs (#1328) · 9d9072a0
  Zhuohan Li authored Oct 16, 2023
```
Co-authored-by: Yunmo Chen <16273544+wanmok@users.noreply.github.com>
```
  9d9072a0
- Implement PagedAttention V2 (#1348) · 928de468
  Woosuk Kwon authored Oct 16, 2023
  
  928de468
11 Oct, 2023 1 commit
- change the timing of sorting logits (#1309) · 91fce82c
  yhlskt23 authored Oct 11, 2023
  
  91fce82c
02 Oct, 2023 2 commits
- TP/quantization/weight loading refactor part 1 - Simplify parallel linear logic (#1181) · ba0bfd40
  Zhuohan Li authored Oct 02, 2023
  
  ba0bfd40
- [Minor] Fix type annotations (#1238) · 84e4e37d
  Woosuk Kwon authored Oct 02, 2023
  
  84e4e37d
28 Sep, 2023 1 commit
- [Mistral] Mistral-7B-v0.1 support (#1196) · bb1ba58f
  Chris Bamford authored Sep 28, 2023
```
Co-authored-by: timlacroix <t@mistral.ai>
```
  bb1ba58f
27 Sep, 2023 1 commit

Support Longchat and RoPE scaling (#555) · 21877b0d

Lily Liu authored Sep 27, 2023


Co-authored-by: Wing Lian <wing.lian@gmail.com>
Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>

21877b0d

26 Sep, 2023 1 commit
- Add comments on RoPE initialization (#1176) · 03ffd0a0
  Woosuk Kwon authored Sep 26, 2023
  
  03ffd0a0
24 Sep, 2023 1 commit
- [FIX] Simplify sampler logic (#1156) · f1878779
  Zhuohan Li authored Sep 23, 2023
  
  f1878779
23 Sep, 2023 1 commit
- [Sampler] Vectorized sampling (simplified) (#1048) · 947b7941
  Zhuohan Li authored Sep 22, 2023
```
Co-authored-by: Antoni Baum <antoni.baum@protonmail.com>
```
  947b7941
16 Sep, 2023 1 commit

Implement AWQ quantization support for LLaMA (#1032) · e3e79e9e

Woosuk Kwon authored Sep 16, 2023


Co-authored-by: Robert Irvine <robert@seamlessml.com>
Co-authored-by: root <rirv938@gmail.com>
Co-authored-by: Casper <casperbh.96@gmail.com>
Co-authored-by: julian-q <julianhquevedo@gmail.com>

e3e79e9e

13 Sep, 2023 1 commit
- [FIX] Minor bug fixes (#1035) · f04908ca
  Zhuohan Li authored Sep 13, 2023
```
* [FIX] Minor bug fixes

* Address review comments
```
  f04908ca
11 Sep, 2023 1 commit
- Use FP32 in RoPE initialization (#1004) · e67b4f2c
  Woosuk Kwon authored Sep 11, 2023
```
Co-authored-by: One <imone@tuta.io>
```
  e67b4f2c
09 Sep, 2023 1 commit
- Fix wrong dtype in PagedAttentionWithALiBi bias (#996) · a62de9ec
  Antoni Baum authored Sep 09, 2023
```
---------
Signed-off-by: Antoni Baum <antoni.baum@protonmail.com>
```
  a62de9ec
08 Sep, 2023 1 commit
- faster startup of vLLM (#982) · 4b5bcf89
  Robert Irvine authored Sep 08, 2023
```
* update

---------
Co-authored-by: Robert Irvine <robert@seamlessml.com>
```
  4b5bcf89
06 Sep, 2023 1 commit
- [BugFix] Implement RoPE for GPT-J (#941) · 320a622e
  Woosuk Kwon authored Sep 06, 2023
  
  320a622e
05 Sep, 2023 1 commit
- Align vLLM's beam search implementation with HF generate (#857) · 002800f0
  Zhuohan Li authored Sep 04, 2023
  
  002800f0
31 Aug, 2023 2 commits
- fix: bug fix when penalties are negative (#913) · e1122233
  Dong-Yong Lee authored Sep 01, 2023
```
Co-authored-by: dongyong-lee <dongyong.lee@navercorp.com>
```
  e1122233
- Improve _prune_hidden_states micro-benchmark (#707) · 28873a27
  Aman Gupta Karmani authored Aug 31, 2023
  
  28873a27
30 Aug, 2023 1 commit
- use flash-attn via xformers (#877) · 75471386
  Aman Gupta Karmani authored Aug 30, 2023
  
  75471386
25 Aug, 2023 1 commit
- Set replacement=True in torch.multinomial (#858) · 94d2f598
  Woosuk Kwon authored Aug 25, 2023
  
  94d2f598
23 Aug, 2023 1 commit
- Fix for breaking changes in xformers 0.0.21 (#834) · 2a4ec908
  Woosuk Kwon authored Aug 23, 2023
  
  2a4ec908
22 Aug, 2023 1 commit
- Implement approximate GELU kernels (#828) · d64bf164
  Woosuk Kwon authored Aug 23, 2023
  
  d64bf164