Commits · 51f5401893f014b2ba131d57f05da8f17c4d4738 · OpenDAS / text-generation-inference

14 Oct, 2024 3 commits

Clarify gated description and quicktour (#2631) · 51f54018
Omar Sanseviero authored Oct 14, 2024
```
Update quicktour.md
```
51f54018

Nicolas Patry authored Oct 14, 2024



* break when there's nothing to read
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>

* Different approach, only listen on stdin when `LOG_LEVEL=debug` (which
is where dropping to a debugger is important).

---------
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>
Co-authored-by: Wang, Yi A <yi.a.wang@intel.com>

3ea82d00

Small fixes for supported models (#2471) · ce28ee88

Omar Sanseviero authored Oct 14, 2024



* Small improvements for docs

* Update _toctree.yml

* Updating the doc (we keep the list actually).

---------
Co-authored-by: Nicolas Patry <patry.nicolas@protonmail.com>

ce28ee88

11 Oct, 2024 1 commit
- Fixing intel Supports windowing. (#2637) · 0c478846
  Nicolas Patry authored Oct 11, 2024
  
  0c478846
10 Oct, 2024 3 commits

Intel ci (#2630) · 3dbdf63e

Nicolas Patry authored Oct 10, 2024

* Intel CI ?

* Let's try non sharded gemma.

* Snapshot rename

* Apparently container can be gone already.

3dbdf63e

Update documentation to most recent stable version of TGI. (#2625) · d912f0bf
vb authored Oct 10, 2024
```
Update to most recent stable version of TGI.
```
d912f0bf

feat: allow tool calling to respond without a tool (#2614) · e36dfaa8

drbh authored Oct 10, 2024



* feat: process token stream before returning to client

* fix: expect content in test

* fix: improve comparison via ruff lint

* fix: return event in all cases

* fix: always send event on error, avoid unwraps, refactor and improve tests

* fix: prefer no_tool over notify_error to improve reponse

* fix: adjust chat input test for no_tool

* fix: adjust test expected content

---------
Co-authored-by: System administrator <root@ip-10-90-0-186.ec2.internal>

e36dfaa8

09 Oct, 2024 2 commits

AMD CI (#2589) · 43f39f68

Nicolas Patry authored Oct 09, 2024

* Only run 1 valid test.

* TRying the tailscale action quickly.

* ?

* bash spaces.

* Remove tailscale.

* More quotes.

* mnt2 ?

* Othername to avoid recursive directories.

* Good old tmate.

* Remove tmate.

* Trying a few things.

* Remove some stuff.

* Sleep ?

* Tmp

* busybox

* Launcher tgi

* Starting hello

* Busybox in python

* No device.

* Removing all variables ?

* A un moment donné.

* Tmp

* Tmp2

* DEvice request, no container name

* No device requests

* Without pytest.

* No pytest.

* from env

* Start with devices

* Attemp #1

* Remove stdin messing

* Only 1 test, no container name

* Raw tgi

* Sending args.

* Show pip freeze.

* Start downloading with token

* Giving HIP devices.

* Mount volume + port forward

* Without pytest.

* No token

* Repeated arguments

* Wrong kwarg.

* On 2 GPUs

* Fallback to single shard CI test.

* Testing

* yaml

* Common cache ?

* Trailing slash ?

* Docker volume split.

* Fix docker volume

* Fixing ?

* ?

* Try no devices ?

* Flash llama on intel CPU ?

* Fix nvidia ?

* Temp deactivate intel, activate nvidia ?

43f39f68

nix: add black and isort to the closure (#2619) · 9ed0c85f

Daniël de Kok authored Oct 09, 2024

To make sure that everything is formatted with the same black version
as CI.

I sometimes use isort for new files to get nicely ordered imports,
so add it as well. Also set the isort configuration to format in a
way that is compatible with black.

9ed0c85f

08 Oct, 2024 4 commits
- CI (2599): Update ToolType input schema (#2601) · 8ad20daf
  drbh authored Oct 08, 2024
```
* Update ToolType input schema

* lint

* fix: run formatter

* fix: allow tool choide to be null

---------
Co-authored-by: Wauplin <lucainp@gmail.com>
```
  8ad20daf
- nix: move back to the tgi-nix main branch (#2620) · 6db3bcb7
  Daniël de Kok authored Oct 08, 2024
  
  6db3bcb7
- Add support for fused MoE Marlin for AWQ (#2616) · 64142489
  Daniël de Kok authored Oct 08, 2024
```
* Add support for fused MoE Marlin for AWQ

This uses the updated MoE Marlin kernels from vLLM.

* Add integration test for AWQ MoE
```
  64142489
- Upgrade minor rust version (Fixes rust build compilation cache) (#2617) · 8b295aa4
  Nicolas Patry authored Oct 08, 2024
```
* Upgrade minor rust version (Fixes rust build compilation cache)

* Black
```
  8b295aa4
07 Oct, 2024 2 commits
- enable mllama in intel platform (#2610) · 57f9685d
  Wang, Yi authored Oct 08, 2024
```
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>
```
  57f9685d
- Fix FP8 KV-cache condition (#2611) · 0da4df4b
  Florian Zimmermeister authored Oct 07, 2024
```
Update kv_cache.py
```
  0da4df4b
04 Oct, 2024 2 commits

Add basic FP8 KV cache support (#2603) · 2358c2bb

Daniël de Kok authored Oct 04, 2024

* Add basic FP8 KV cache support

This change adds rudimentary FP8 KV cache support. The support is
enabled by passing `--kv-cache-dtype fp8_e5m2` to the launcher. Doing so
uses this type for the KV cache. However support is still limited:

* Only the `fp8_e5m2` type is supported.
* The KV cache layout is the same as `float16`/`bfloat16` (HND).
* The FP8 KV cache is only supported for FlashInfer.
* Loading of scales is not yet supported.

* Fix Cargo.toml

2358c2bb

nix: example of local package overrides during development (#2607) · 68103079
Daniël de Kok authored Oct 04, 2024

68103079

03 Oct, 2024 2 commits
- Revert "Unroll notify error into generate response" (#2605) · 3011639f
  drbh authored Oct 03, 2024
```
Revert "Unroll notify error into generate response (#2597)"

This reverts commit d22b0c1f.
```
  3011639f
- New release 2.3.1 (#2604) · f6e2f05b
  Nicolas Patry authored Oct 03, 2024
```
* New release 2.3.1

* Update doc number
```
  f6e2f05b
02 Oct, 2024 4 commits

Unroll notify error into generate response (#2597) · d22b0c1f

drbh authored Oct 02, 2024

* feat: unroll notify_error if no tool is choosen

* fix: expect simple message when no tool is selected

* fix: improve test to avoid notify_error

* fix: improve docs and indicate change in expected response

* fix: adjust linting in test file

d22b0c1f

CI (2592): Allow LoRA adapter revision in server launcher (#2602) · 23354595

drbh authored Oct 02, 2024



allow revision for lora adapters from launcher
Co-authored-by: Sida <sida@kulamind.com>
Co-authored-by: teamclouday <teamclouday@gmail.com>

23354595

Max token capacity metric (#2595) · 0204946d

Nicolas Patry authored Oct 02, 2024



* adding max_token_capacity_metric

* added tgi to name of metric

* Adding max capacity metric.

* Add description for the metrics

---------
Co-authored-by: Edwinhr716 <Edandres249@gmail.com>

0204946d

Mllama flash version (#2585) · d18ed5cf

Nicolas Patry authored Oct 02, 2024

* Working loading state.

* Preprocessing.

* Working state ? (Broke idefics1 temporarily).

* Cleaner condition.

* Fix idefics.

* Updating config, removing TODO

* Mllama

* Ugrade transformers 4.45

* Flashing mllama.

* Starting to get there.

* Working state.

* Integrations tests for mllama (cutting to 10 tokens because there seems'
to be instability after (meaning size of the batch matters.

* Updating model link.

* Earlier assert.

* Fix vlm ?

* remove log.

* Force ignore all images but last.

* Default dtype bfloat16.

* Update integration test after switch to bf16.

* Remove dead code.

* Removed dead code.

* Upgrade the flake to latest transformers/tokenizers

* Move to hf tgi-nix

* Upgrade to 0.5.0

d18ed5cf

01 Oct, 2024 1 commit

nix: experimental support for building a Docker container (#2470) · 584b4d7a

Daniël de Kok authored Oct 01, 2024



* nix: experimental support for building a Docker image

Run using something like:

```
docker run \
  --device nvidia.com/gpu=all \
  -it --rm -p 8080:80 \
  -v $PWD/data:/data \
  -v $PWD/tmp:/tmp \
  tgi-docker:latest \
  --model-id <model_id>
```

* Example of building the Docker image using Nix inside Docker

* Stream to make the builder image smaller

This avoids storing a Docker image tarball in the image. Instead,
stream the layers while doing `docker run`.

* Don't spam journalctl on Linux

* Other dockerfile.

---------
Co-authored-by: Nicolas Patry <patry.nicolas@protonmail.com>

584b4d7a

30 Sep, 2024 7 commits

MoE Marlin: support `desc_act` for `groupsize != -1` (#2590) · 1c84a30f
Daniël de Kok authored Sep 30, 2024
```
This change uses the updated Marlin MoE kernel from vLLM to support
MoE with activation sorting and groups.
```
1c84a30f
Move flake back to tgi-nix `main` (#2586) · d1f257ac
Daniël de Kok authored Sep 30, 2024

d1f257ac

feat: support phi3.5 moe (#2479) · 93a7042d

drbh authored Sep 30, 2024



* feat: support phi3.5 moe model loading

* fix: prefer llama base model and improve rotary logic

* feat: return reasonable generation and add integration test

* fix: run lint and update docs

* fix: rerun lint for openapi docs

* fix: prefer do_sample false unless temp is set by user, and update chat tests

* fix: small typo adjustments

* fix: consolidate long rope paths

* fix: revert greedy by default and test changes

* Vendor configuration so that we don't have to `trust_remote_code`

* Use SparseMoELayer

* Add support for dense MoE

* Some type annotations

* Add the usual model tests

* Ruff.

---------
Co-authored-by: Daniël de Kok <me@danieldk.eu>
Co-authored-by: Nicolas Patry <patry.nicolas@protonmail.com>

93a7042d

Add support for GPTQ-quantized MoE models using MoE Marlin (#2557) · 90a1d04a

Daniël de Kok authored Sep 30, 2024

This change add support for MoE models that use GPTQ quantization.
Currently only models with the following properties are supported:

- No `desc_act` with tensor parallelism, unless `group_size=-1`.
- No asymmetric quantization.
- No AWQ.

90a1d04a

Update ROCM libs and improvements (#2579) · f9e561ec

Mohit Sharma authored Sep 30, 2024

* style

* update torch

* ix issues

* fix clone

* revert mkl

* added custom PA

* style

* fix style

* style

* hide env vart

* fix mixtral model

* add skinny kernel and merge fixes

* fixed style

* fix issue for sliding window models

* addressed review comments

* fix import

* improved error messag

* updated default value

* remove import

* fix imports after rebase

* float16 dep

* improve dockerfile

* cleaned dockerfile

f9e561ec

Update architecture.md (#2577) · e790cfc0
Ikram Ul Haq authored Sep 30, 2024

e790cfc0

Remove compute capability lazy cell (#2580) · afc7ded8

Daniël de Kok authored Sep 30, 2024

Remove compute capability lock

We are only calling the `get_cuda_capability` function once, so avoiding
the cost of multiple calls is not really necessary yet.

afc7ded8

28 Sep, 2024 1 commit
- flashinfer: pass window size and dtype (#2574) · 1028996f
  Daniël de Kok authored Sep 28, 2024
  
  1028996f
27 Sep, 2024 1 commit

Improve support for GPUs with capability < 8 (#2575) · 5b6b74e2

Daniël de Kok authored Sep 27, 2024

* Improve support for GPUs with capability < 8

- For models that cannot use flashinfer, use flash-attn v1 + paged
  attention for models with a compute capability older than 8.
- Disable prefix caching when using paged attention.
- When using flash-attn v1, pass the key/value, rather than the
  cache, since v1 cannot use block tables.

* nix: add flash-attn-v1 to the server environment

* Move disabling prefix caching into the block of exceptions

* Capability as `usize`s

5b6b74e2

26 Sep, 2024 2 commits
- Fix build with `--features google` (#2566) · 0aa66d69
  Alvaro Bartolome authored Sep 26, 2024
```
* Fix `cargo build --features google`

* Add `cargo test --features google`
```
  0aa66d69
- Add LoRA adapters support for Gemma2 (#2567) · 0b7df771
  Alvaro Bartolome authored Sep 26, 2024
```
* Add LoRA adapters support for Gemma2

* Make `black` formatting happy
```
  0b7df771
24 Sep, 2024 5 commits

remove LORA_ADAPTERS_PATH (#2563) · 7efcb5e0
Nicholas Broad authored Sep 24, 2024
```
specify how to call local adapters
```
7efcb5e0
More tensor cores. (#2558) · dd8691b7
Nicolas Patry authored Sep 24, 2024
```
* More tensor cores.

* Fixing the logic.

* Gemma is modified by this.
```
dd8691b7

Cleanup Vertex + Chat (#2553) · c032280b

Nicolas Patry authored Sep 24, 2024

* Cleanup Vertex + Chat

* logprobs defaults to false.

* Parameters are optional

* Fix  docs.

* Changing back this logprobs default.

* Fixup doc.

* Let's debug that.

* Not unstable.

* Updating Cargo ?

* Wat?

* Dummy change.

* Trying some other install.

* Trying smething.

* Revert everything.

* Update Cargo lock.

* Fixing the pre-commit after rebase.

c032280b

Hotfixing main. (#2562) · 75c8c54a
Nicolas Patry authored Sep 24, 2024

75c8c54a

Adding note for private models in quick-tour document (#2548) · e6d29656

Aritra Roy Gosthipaty authored Sep 24, 2024



* chore: adding note for private models in quicktour doc

* Update docs/source/quicktour.md
Co-authored-by: Omar Sanseviero <osanseviero@gmail.com>

* Update docs/source/quicktour.md
Co-authored-by: vb <vaibhavs10@gmail.com>

* Update docs/source/quicktour.md
Co-authored-by: vb <vaibhavs10@gmail.com>

---------
Co-authored-by: Omar Sanseviero <osanseviero@gmail.com>
Co-authored-by: vb <vaibhavs10@gmail.com>

e6d29656