Commits · 4c976fb4064f95b6604745b81c91c6b7bbd20072 · OpenDAS / text-generation-inference

08 Jul, 2024 5 commits
- Fix nccl regression on PyTorch 2.3 upgrade (#2099) · 4c50b6d0
  fxmarty authored Jul 08, 2024
```
* fix nccl issue

* add note in dockerfile

* use v2.22.3 that also fixes @samsamoa's repro

* poetry actually can't handle the conflict between torch and nccl

* set LD_PRELOAD
```
  4c50b6d0
- Falcon/DBRX: get correct number of key-value heads (#2205) · 5c7c9f13
  Daniël de Kok authored Jul 08, 2024
  
  5c7c9f13
- Fix incorrect cache allocation with multi-query (#2203) · 153fcf77
  Daniël de Kok authored Jul 08, 2024
```
We wouldn't allocate any memory in multi-query (1 KV head). Fixes
Starcoder et al.
```
  153fcf77
- hotfix: Fix number of KV heads (#2202) · cce475a9
  Daniël de Kok authored Jul 08, 2024
```
Fix number of KV heads
```
  cce475a9
- fix dbrx & opt model prefix bug (#2201) · 521d0d99
  icyboy™ authored Jul 08, 2024
```
* Update idefics_causal_lm.py

Fix syntax issues

* fix dbrx & opt model prefix bug
```
  521d0d99
05 Jul, 2024 5 commits

Consistently take `prefix` in model constructors (#2191) · 05c094fc
Daniël de Kok authored Jul 05, 2024
```
* Consistently take `prefix` in model constructors

* Release test check fix

* Misc refactor-related fixes
```
05c094fc
Fix Starcoder2 after refactor (#2189) · b67d4633
Daniël de Kok authored Jul 05, 2024

b67d4633
Hotfixing after refactor. · 853d4eb9
Nicolas Patry authored Jul 05, 2024

853d4eb9

Refactor dead code - Removing all `flash_xxx.py` files. (#2166) · fb2f74e2

Nicolas Patry authored Jul 05, 2024

* Refactor dead code.

* First working step.

* Remove a lot of duplicated code.

* More dead code.

* More cleanup.

* Fix Santacoder test.

* Fixing the simple tests.

* Fixing sharding.

* Fixes for VLM.

* Fixing santacoder (num_kv_heads hardcoded).

* Removing more dead code.

* Fixing `config.n_head`.

* Stopping earlier because of `<end_of_utterance>` in idefics2.

* Addresses comments.

* Removing the dead code.

* Fuse back mistral into FlashCausalLM.

* Finish removal.

* Fixing docs + causal_lm `batch_class`.

* Fixing docs + causal.lm.

* Add default to Gemma Causality.

* Default value for gemma/gemma2.

* Wrong default.

fb2f74e2

Adding "longrope" for Phi-3 (#2172) (#2179) · c6bcadf8
Aaron Mihalik authored Jul 05, 2024
```
Adding "longrope" for phi-3
```
c6bcadf8

02 Jul, 2024 5 commits
- Hotfixing qwen2 and starcoder2 (which also get clamping). (#2167) · 0759ec49
  Nicolas Patry authored Jul 02, 2024
  
  0759ec49
- Fixing rocm. (#2164) · dea9c0dc
  Nicolas Patry authored Jul 02, 2024
  
  dea9c0dc
- fix: use the base layers weight in mistral rocm (#2155) · b966bc0d
  drbh authored Jul 02, 2024
  
  b966bc0d
- fix FlashDecoding change's regression in intel platform (#2161) · 5d97e0c4
  Wang, Yi authored Jul 02, 2024
```
install triton because GPTQParams needs it.
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>
```
  5d97e0c4
- Fixing graph capture for flash decoding. (#2163) · 022f6515
  Nicolas Patry authored Jul 02, 2024
  
  022f6515
01 Jul, 2024 6 commits

[Major Change][Undecided yet] Move to FlashDecoding instead of PagedAttention kernel. (#1940) · 4327210e

Nicolas Patry authored Jul 01, 2024

* Using flash decoding

Conditional flashdecoding.

Fix max_q.

Working kvcache

Working version with flash decoding.

Make it work for mistral.

Fix after rebase..

Less intrusive.

REvert changes in modeling.

Speedup flashdecoding.

HHachweew
Hack to make other models work.

Fixing non flash decoding llama path.

Router logic knows about page size.

Missing 2 models.

Missing cohere.

Fixing cohere flash decoding.

Revamped all this architecture.

Fix cohere.

Fixing falcon.

Enabling custom block size schedule.

Update router/src/infer.rs

Not sending preallocated output.

* Making it work on non flash decoding.

* Fix Cohere.

* Fix non decoding paths.

* Rebased.

* No need for cache_manager anymore.

* Update?

* "ipex" -> "cpu"

* These do not belong.

* Factoring cu_seqlen_qk for better abstracting over every model.

* Fixing non flash tests/imports.

* Changing return everywhere.

* Update mistral past.

* Fixing Mi{s,x}tral (non functional in Flash Decoding mode though).

* Fixup mistral clamping (had issues with cuda graphs).

* No need to recreate anything actually.

4327210e

Fixing baichuan override. (#2158) · 4f55f158
Nicolas Patry authored Jul 01, 2024

4f55f158

refine get xpu free memory/enable Qwen2/gemma2/gemma/phi in intel platform (#2132) · 5da4cfab

Wang, Yi authored Jul 01, 2024



* refine get xpu free memory
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>

* enable qwen2 in xpu
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>

* enable gemma/gemma2/phi in intel platform
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>

---------
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>

5da4cfab

fix AttributeError: 'MixtralLayer' object has no attribute 'mlp' (#2123) · 9d0ca503
icyboy™ authored Jul 01, 2024
```
https://github.com/huggingface/text-generation-inference/issues/2122
```
9d0ca503

Use GPTQ-Marlin for supported GPTQ configurations (#2111) · 2ce80194

Daniël de Kok authored Jul 01, 2024

GPTQ-Marlin is currently the best-performing kernel for GPTQ models. So
let's use it by default if the kernels are installed, the GPU supports
it, and the kernels support the configuration.

For models generated by `text-generation-server quantize`, use
`sym=False`. This subcommand symmetric quantization since the beginning
and incorrectly reporting the model to be symmetric will use
GPTQ-Marlin (which does not support asymmetric quantization).

2ce80194

fix: use weights from base_layer (#2141) · 25f57e2e
drbh authored Jul 01, 2024

25f57e2e

27 Jun, 2024 2 commits

Fixing gemma2. (#2135) · 3ea8259a
Nicolas Patry authored Jun 27, 2024
```
* Fixing gemma2.

* Adding new model.
```
3ea8259a

Idefics2: sync added image tokens with transformers (#2080) · dd2d91b0

Daniël de Kok authored Jun 27, 2024

Before this change, the number of reserved image tokens was not the
same as the number of images. Fixes #2029.

While at it, also remove all the image token handling duplication
in `prepare_input`.

dd2d91b0

25 Jun, 2024 9 commits

Add support for Marlin 2:4 sparsity (#2102) · f1f98e36

Daniël de Kok authored Jun 25, 2024

This change adds support for 2:4 sparsity when using Marlin
quantization. The 2:4 kernel is used when:

* The quantizer is `marlin`;
* the quantizer checkpoint format is `marlin_24`.

Fixes #2098.

f1f98e36

Support AWQ quantization with bias (#2117) · 14980df2

Daniël de Kok authored Jun 25, 2024

When the AWQ quantizer was used with a layer that uses a bias,
the bias tensor was not correctly passed/used. Instead, the
value `true`/`1.0` was added to the linear transformation.

Correctly pass through the bias when it is not `None`.

Fixes #2106.

14980df2

Enable multiple LoRa adapters (#2010) · 04e1af94

drbh authored Jun 25, 2024



* feat: first draft load multiple lora

* feat: load weights within layer and refactor lora pass

* fix: refactor and reduce lora math

* feat: baseline impl single request multi lora support

* feat: prefer lorax implementation and port loading logic

* fix: prefer adapter_data and refactors

* feat: perfer loraxs custom punica kernels and add mlp loras

* fix: adjust batch for bgmv

* fix: adjust adapter_segments logic when in batch

* fix: refactor and move changes to v3 proto

* fix: pass model_id for all flash causal lms

* fix: pass model_id for all causal and seq2seq lms

* fix: add model_id to model test

* feat: add lora support to mistral and refactors

* feat: prefer model id in request

* fix: include rust code for adapter id

* feat: bump launcher and add new lora docs

* feat: support base model generation and refactors

* fix: rename doc to retry ci build

* feat: support if vlm models

* fix: add adapter_data param and avoid missing layers

* fix: add adapter_data param to phi and neox

* fix: update all models forwards to include adapter_data

* fix: add model_id to IdeficsCausalLM

* Update lora.md

Fixed a typo

* Update lora.md

Fixing spam image

* fix: add lora kernel to dockerfile, support running without kernels and refactors

* fix: avoid dockerfile conflict

* fix: refactors and adjust flash llama lora logic

* fix: skip llama test due to CI issue (temp)

* fix: skip llama test CI (temp) 2

* fix: revert skips and prefer updated ci token for tests

* fix: refactors and helpful comments

* fix: add noop in TensorParallelAdapterRowLinear too

* fix: refactor and move shard_lora_weights logic

* fix: exit early if no adapter_data

---------
Co-authored-by: Derek <datavistics@gmail.com>

04e1af94

fix cpu and xpu issue (#2116) · e563983d
Wang, Yi authored Jun 25, 2024
```
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>
```
e563983d

Removing IPEX_AVAIL. (#2115) · 9e2fdf57

Nicolas Patry authored Jun 25, 2024

* Removing IPEX_AVAIL.

Chose to unify CPU and XPU under `ipex`. Most code is exactly similar
except for a very few spots.

The biggest number of spots is the kv-cache layout and the flash_xxx.py
files.
Since those files should be removed soon and factored away, we should
not need them.

* Forgot a few places.

* Unrelated change.

* Fixing HF_TOKEN.

* HF_TOKEN

9e2fdf57

feat: add simple tests for weights (#2092) · 3f3b7ffd

drbh authored Jun 25, 2024

* feat: add simple tests for weights

* fix: adjust types and add tests

* fix: adjust so all tests pass

* feat: improve weight tests

* fix: add missing tests and renames

* fix: tweak shapes

3f3b7ffd

Cpu tgi (#1936) · b64c70c9

Wang, Yi authored Jun 25, 2024



* add CPU tgi support
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>

* ipex distributed ops support
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>

---------
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>
Co-authored-by: Funtowicz Morgan <mfuntowicz@users.noreply.github.com>

b64c70c9

use xpu-smi to dump used memory (#2047) · 83634dc1

Wang, Yi authored Jun 25, 2024



* use xpu-smi to dump used memory
xpu use "ZE_AFFINITY_MASK" to control card, usage is like CUDA_VISIBLE_DEVICES
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>

* Update server/text_generation_server/utils/import_utils.py
Co-authored-by: Daniël de Kok <me@github.danieldk.eu>

---------
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>
Co-authored-by: Daniël de Kok <me@github.danieldk.eu>

83634dc1

Add OTLP Service Name Environment Variable (#2076) · 1869ee2f

KevinDuffy94 authored Jun 25, 2024

* Adding Service Name Environment variable for https://github.com/huggingface/text-generation-inference/issues/2069

* Update Docs

* Update README.md

* Update Launcher Docs

* Update Launcher Docs
Removing Option

1869ee2f

21 Jun, 2024 2 commits
- feat: sort cuda graphs in descending order (#2104) · 811a9381
  drbh authored Jun 21, 2024
  
  811a9381
- Fix `text-generation-server quantize` (#2103) · 197c47a3
  Daniël de Kok authored Jun 21, 2024
```
The subcommand did not work due to some broken imports.
```
  197c47a3
20 Jun, 2024 2 commits

Factor out sharding of packed tensors (#2059) · bcb3faa1

Daniël de Kok authored Jun 20, 2024

For Phi-3-Small I need to shard a packed QKV bias tensor, for which
I implemented the `Weights.get_packed_sharded` method. However, this
method can also replace the `Weights._get_qweight` method and the
custom sharding code from `Weights.get_weights_col_packed`.

bcb3faa1

Support exl2-quantized Qwen2 models (#2085) · f5a98375
Daniël de Kok authored Jun 20, 2024
```
Fixes #2081.
```
f5a98375

17 Jun, 2024 2 commits

Set maximum grpc message receive size to 2GiB (#2075) · c8c7ccd3

Daniël de Kok authored Jun 17, 2024

* Set maximum grpc message receive size to 2GiB

The previous default was 4MiB, which doesn't really work well for
multi-modal models.

* Update to Rust 1.79.0

* Fixup formatting to make PR pass

c8c7ccd3

Support different image sizes in prefill in VLMs (#2065) · e9037708

Daniël de Kok authored Jun 17, 2024

When a batch contained images if different sizes during prefill, the
server would fail (see e.g. #2056). Images were processed separately and
then concatenated. However, this can fail for images with different sizes.

Fix this by preprocessing all images in the batch together, so that the
image processor can ensure that all image tensors have compatible sizes.

e9037708

14 Jun, 2024 2 commits

Update the link for qwen2 (#2068) · 96b7b40c

Tiezhen WANG authored Jun 14, 2024



* Update the link for qwen2

* Fix Qwen2 model URL in model table

* Fix too eager staging

---------
Co-authored-by: Daniël de Kok <me@danieldk.eu>

96b7b40c

Add support for GPTQ Marlin (#2052) · 093a27c5

Daniël de Kok authored Jun 14, 2024

Add support for GPTQ Marlin kernels

GPTQ Marlin extends the Marlin kernels to support common GPTQ
configurations:

- bits: 4 or 8
- groupsize: -1, 32, 64, or 128
- desc_act: true/false

Using the GPTQ Marlin kernels requires repacking the parameters in the
Marlin quantizer format.

The kernels were contributed by Neural Magic to VLLM. We vendor them
here for convenience.

093a27c5