Commits · cf04a43fb1f9aa08f8095e99b9ef04cd768af302 · OpenDAS / text-generation-inference

15 Oct, 2024 1 commit
- Fixing linters. (#2650) · cf04a43f
  Nicolas Patry authored Oct 15, 2024
  
  cf04a43f
14 Oct, 2024 5 commits

feat: enable pytorch xpu support for non-attention models (#2561) · 58848cb4

Dmitry Rogozhkin authored Oct 14, 2024



XPU backend is available natively (without IPEX) in pytorch starting
from pytorch 2.4. This commit extends TGI to cover the case when user
has XPU support thru pytorch 2.4, but does not have IPEX installed.
Models which don't require attention can work. For attention required
models more work is needed to provide attention implementation.

Tested with the following models:
* teknium/OpenHermes-2.5-Mistral-7B
* bigscience/bloom-560m
* google/gemma-7b
* google/flan-t5-xxl
Signed-off-by: Dmitry Rogozhkin <dmitry.v.rogozhkin@intel.com>

58848cb4

update ipex to fix incorrect output of mllama in cpu (#2640) · 7a82ddcb
Wang, Yi authored Oct 14, 2024
```
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>
```
7a82ddcb
Clarify gated description and quicktour (#2631) · 51f54018
Omar Sanseviero authored Oct 14, 2024
```
Update quicktour.md
```
51f54018

Cpu perf (#2596) · 3ea82d00

Nicolas Patry authored Oct 14, 2024



* break when there's nothing to read
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>

* Different approach, only listen on stdin when `LOG_LEVEL=debug` (which
is where dropping to a debugger is important).

---------
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>
Co-authored-by: Wang, Yi A <yi.a.wang@intel.com>

3ea82d00

Small fixes for supported models (#2471) · ce28ee88

Omar Sanseviero authored Oct 14, 2024



* Small improvements for docs

* Update _toctree.yml

* Updating the doc (we keep the list actually).

---------
Co-authored-by: Nicolas Patry <patry.nicolas@protonmail.com>

ce28ee88

11 Oct, 2024 1 commit
- Fixing intel Supports windowing. (#2637) · 0c478846
  Nicolas Patry authored Oct 11, 2024
  
  0c478846
10 Oct, 2024 3 commits

Intel ci (#2630) · 3dbdf63e

Nicolas Patry authored Oct 10, 2024

* Intel CI ?

* Let's try non sharded gemma.

* Snapshot rename

* Apparently container can be gone already.

3dbdf63e

Update documentation to most recent stable version of TGI. (#2625) · d912f0bf
vb authored Oct 10, 2024
```
Update to most recent stable version of TGI.
```
d912f0bf

feat: allow tool calling to respond without a tool (#2614) · e36dfaa8

drbh authored Oct 10, 2024



* feat: process token stream before returning to client

* fix: expect content in test

* fix: improve comparison via ruff lint

* fix: return event in all cases

* fix: always send event on error, avoid unwraps, refactor and improve tests

* fix: prefer no_tool over notify_error to improve reponse

* fix: adjust chat input test for no_tool

* fix: adjust test expected content

---------
Co-authored-by: System administrator <root@ip-10-90-0-186.ec2.internal>

e36dfaa8

09 Oct, 2024 2 commits

AMD CI (#2589) · 43f39f68

Nicolas Patry authored Oct 09, 2024

* Only run 1 valid test.

* TRying the tailscale action quickly.

* ?

* bash spaces.

* Remove tailscale.

* More quotes.

* mnt2 ?

* Othername to avoid recursive directories.

* Good old tmate.

* Remove tmate.

* Trying a few things.

* Remove some stuff.

* Sleep ?

* Tmp

* busybox

* Launcher tgi

* Starting hello

* Busybox in python

* No device.

* Removing all variables ?

* A un moment donné.

* Tmp

* Tmp2

* DEvice request, no container name

* No device requests

* Without pytest.

* No pytest.

* from env

* Start with devices

* Attemp #1

* Remove stdin messing

* Only 1 test, no container name

* Raw tgi

* Sending args.

* Show pip freeze.

* Start downloading with token

* Giving HIP devices.

* Mount volume + port forward

* Without pytest.

* No token

* Repeated arguments

* Wrong kwarg.

* On 2 GPUs

* Fallback to single shard CI test.

* Testing

* yaml

* Common cache ?

* Trailing slash ?

* Docker volume split.

* Fix docker volume

* Fixing ?

* ?

* Try no devices ?

* Flash llama on intel CPU ?

* Fix nvidia ?

* Temp deactivate intel, activate nvidia ?

43f39f68

nix: add black and isort to the closure (#2619) · 9ed0c85f

Daniël de Kok authored Oct 09, 2024

To make sure that everything is formatted with the same black version
as CI.

I sometimes use isort for new files to get nicely ordered imports,
so add it as well. Also set the isort configuration to format in a
way that is compatible with black.

9ed0c85f

08 Oct, 2024 4 commits
- CI (2599): Update ToolType input schema (#2601) · 8ad20daf
  drbh authored Oct 08, 2024
```
* Update ToolType input schema

* lint

* fix: run formatter

* fix: allow tool choide to be null

---------
Co-authored-by: Wauplin <lucainp@gmail.com>
```
  8ad20daf
- nix: move back to the tgi-nix main branch (#2620) · 6db3bcb7
  Daniël de Kok authored Oct 08, 2024
  
  6db3bcb7
- Add support for fused MoE Marlin for AWQ (#2616) · 64142489
  Daniël de Kok authored Oct 08, 2024
```
* Add support for fused MoE Marlin for AWQ

This uses the updated MoE Marlin kernels from vLLM.

* Add integration test for AWQ MoE
```
  64142489
- Upgrade minor rust version (Fixes rust build compilation cache) (#2617) · 8b295aa4
  Nicolas Patry authored Oct 08, 2024
```
* Upgrade minor rust version (Fixes rust build compilation cache)

* Black
```
  8b295aa4
07 Oct, 2024 2 commits
- enable mllama in intel platform (#2610) · 57f9685d
  Wang, Yi authored Oct 08, 2024
```
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>
```
  57f9685d
- Fix FP8 KV-cache condition (#2611) · 0da4df4b
  Florian Zimmermeister authored Oct 07, 2024
```
Update kv_cache.py
```
  0da4df4b
04 Oct, 2024 2 commits

Add basic FP8 KV cache support (#2603) · 2358c2bb

Daniël de Kok authored Oct 04, 2024

* Add basic FP8 KV cache support

This change adds rudimentary FP8 KV cache support. The support is
enabled by passing `--kv-cache-dtype fp8_e5m2` to the launcher. Doing so
uses this type for the KV cache. However support is still limited:

* Only the `fp8_e5m2` type is supported.
* The KV cache layout is the same as `float16`/`bfloat16` (HND).
* The FP8 KV cache is only supported for FlashInfer.
* Loading of scales is not yet supported.

* Fix Cargo.toml

2358c2bb

nix: example of local package overrides during development (#2607) · 68103079
Daniël de Kok authored Oct 04, 2024

68103079

03 Oct, 2024 2 commits
- Revert "Unroll notify error into generate response" (#2605) · 3011639f
  drbh authored Oct 03, 2024
```
Revert "Unroll notify error into generate response (#2597)"

This reverts commit d22b0c1f.
```
  3011639f
- New release 2.3.1 (#2604) · f6e2f05b
  Nicolas Patry authored Oct 03, 2024
```
* New release 2.3.1

* Update doc number
```
  f6e2f05b
02 Oct, 2024 4 commits

Unroll notify error into generate response (#2597) · d22b0c1f

drbh authored Oct 02, 2024

* feat: unroll notify_error if no tool is choosen

* fix: expect simple message when no tool is selected

* fix: improve test to avoid notify_error

* fix: improve docs and indicate change in expected response

* fix: adjust linting in test file

d22b0c1f

CI (2592): Allow LoRA adapter revision in server launcher (#2602) · 23354595

drbh authored Oct 02, 2024



allow revision for lora adapters from launcher
Co-authored-by: Sida <sida@kulamind.com>
Co-authored-by: teamclouday <teamclouday@gmail.com>

23354595

Max token capacity metric (#2595) · 0204946d

Nicolas Patry authored Oct 02, 2024



* adding max_token_capacity_metric

* added tgi to name of metric

* Adding max capacity metric.

* Add description for the metrics

---------
Co-authored-by: Edwinhr716 <Edandres249@gmail.com>

0204946d

Mllama flash version (#2585) · d18ed5cf

Nicolas Patry authored Oct 02, 2024

* Working loading state.

* Preprocessing.

* Working state ? (Broke idefics1 temporarily).

* Cleaner condition.

* Fix idefics.

* Updating config, removing TODO

* Mllama

* Ugrade transformers 4.45

* Flashing mllama.

* Starting to get there.

* Working state.

* Integrations tests for mllama (cutting to 10 tokens because there seems'
to be instability after (meaning size of the batch matters.

* Updating model link.

* Earlier assert.

* Fix vlm ?

* remove log.

* Force ignore all images but last.

* Default dtype bfloat16.

* Update integration test after switch to bf16.

* Remove dead code.

* Removed dead code.

* Upgrade the flake to latest transformers/tokenizers

* Move to hf tgi-nix

* Upgrade to 0.5.0

d18ed5cf

01 Oct, 2024 1 commit

nix: experimental support for building a Docker container (#2470) · 584b4d7a

Daniël de Kok authored Oct 01, 2024



* nix: experimental support for building a Docker image

Run using something like:

```
docker run \
  --device nvidia.com/gpu=all \
  -it --rm -p 8080:80 \
  -v $PWD/data:/data \
  -v $PWD/tmp:/tmp \
  tgi-docker:latest \
  --model-id <model_id>
```

* Example of building the Docker image using Nix inside Docker

* Stream to make the builder image smaller

This avoids storing a Docker image tarball in the image. Instead,
stream the layers while doing `docker run`.

* Don't spam journalctl on Linux

* Other dockerfile.

---------
Co-authored-by: Nicolas Patry <patry.nicolas@protonmail.com>

584b4d7a

30 Sep, 2024 7 commits

MoE Marlin: support `desc_act` for `groupsize != -1` (#2590) · 1c84a30f
Daniël de Kok authored Sep 30, 2024
```
This change uses the updated Marlin MoE kernel from vLLM to support
MoE with activation sorting and groups.
```
1c84a30f
Move flake back to tgi-nix `main` (#2586) · d1f257ac
Daniël de Kok authored Sep 30, 2024

d1f257ac

feat: support phi3.5 moe (#2479) · 93a7042d

drbh authored Sep 30, 2024



* feat: support phi3.5 moe model loading

* fix: prefer llama base model and improve rotary logic

* feat: return reasonable generation and add integration test

* fix: run lint and update docs

* fix: rerun lint for openapi docs

* fix: prefer do_sample false unless temp is set by user, and update chat tests

* fix: small typo adjustments

* fix: consolidate long rope paths

* fix: revert greedy by default and test changes

* Vendor configuration so that we don't have to `trust_remote_code`

* Use SparseMoELayer

* Add support for dense MoE

* Some type annotations

* Add the usual model tests

* Ruff.

---------
Co-authored-by: Daniël de Kok <me@danieldk.eu>
Co-authored-by: Nicolas Patry <patry.nicolas@protonmail.com>

93a7042d

Add support for GPTQ-quantized MoE models using MoE Marlin (#2557) · 90a1d04a

Daniël de Kok authored Sep 30, 2024

This change add support for MoE models that use GPTQ quantization.
Currently only models with the following properties are supported:

- No `desc_act` with tensor parallelism, unless `group_size=-1`.
- No asymmetric quantization.
- No AWQ.

90a1d04a

Update ROCM libs and improvements (#2579) · f9e561ec

Mohit Sharma authored Sep 30, 2024

* style

* update torch

* ix issues

* fix clone

* revert mkl

* added custom PA

* style

* fix style

* style

* hide env vart

* fix mixtral model

* add skinny kernel and merge fixes

* fixed style

* fix issue for sliding window models

* addressed review comments

* fix import

* improved error messag

* updated default value

* remove import

* fix imports after rebase

* float16 dep

* improve dockerfile

* cleaned dockerfile

f9e561ec

Update architecture.md (#2577) · e790cfc0
Ikram Ul Haq authored Sep 30, 2024

e790cfc0

Remove compute capability lazy cell (#2580) · afc7ded8

Daniël de Kok authored Sep 30, 2024

Remove compute capability lock

We are only calling the `get_cuda_capability` function once, so avoiding
the cost of multiple calls is not really necessary yet.

afc7ded8

28 Sep, 2024 1 commit
- flashinfer: pass window size and dtype (#2574) · 1028996f
  Daniël de Kok authored Sep 28, 2024
  
  1028996f
27 Sep, 2024 1 commit

Improve support for GPUs with capability < 8 (#2575) · 5b6b74e2

Daniël de Kok authored Sep 27, 2024

* Improve support for GPUs with capability < 8

- For models that cannot use flashinfer, use flash-attn v1 + paged
  attention for models with a compute capability older than 8.
- Disable prefix caching when using paged attention.
- When using flash-attn v1, pass the key/value, rather than the
  cache, since v1 cannot use block tables.

* nix: add flash-attn-v1 to the server environment

* Move disabling prefix caching into the block of exceptions

* Capability as `usize`s

5b6b74e2

26 Sep, 2024 2 commits
- Fix build with `--features google` (#2566) · 0aa66d69
  Alvaro Bartolome authored Sep 26, 2024
```
* Fix `cargo build --features google`

* Add `cargo test --features google`
```
  0aa66d69
- Add LoRA adapters support for Gemma2 (#2567) · 0b7df771
  Alvaro Bartolome authored Sep 26, 2024
```
* Add LoRA adapters support for Gemma2

* Make `black` formatting happy
```
  0b7df771
24 Sep, 2024 2 commits
- remove LORA_ADAPTERS_PATH (#2563) · 7efcb5e0
  Nicholas Broad authored Sep 24, 2024
```
specify how to call local adapters
```
  7efcb5e0
- More tensor cores. (#2558) · dd8691b7
  Nicolas Patry authored Sep 24, 2024
```
* More tensor cores.

* Fixing the logic.

* Gemma is modified by this.
```
  dd8691b7