Commits · 46a5a7e73e8b1adde4a279c7d68818ed9e17f607 · OpenDAS / text-generation-inference

"ml/vscode:/vscode.git/clone" did not exist on "1108d8b34e43e968812eded0ccda73503ccad77d"

20 Nov, 2024 1 commit
- Add support for wNa16 int 2:4 compressed-tensors checkpoints (#2758) · 46a5a7e7
  Daniël de Kok authored Nov 20, 2024
```
This change adds support for wNa16 int checkpoints with 2:4 sparsity
using Marlin 2:4 kernels.
```
  46a5a7e7
10 Nov, 2024 1 commit

Add initial support for compressed-tensors checkpoints (#2732) · a7850008

Daniël de Kok authored Nov 10, 2024

compressed-tensors is a safetensors extension for sparse, quantized
tensors. The format is more powerful than earlier AWQ/GPTQ/FP8
quantization, because

- Different quantizer configurations can be used for different targets.
- The format can specify input/output quantizers in addition to weight
  quantizers.
- Configurable exclusions for quantization.

This change adds a dependency on the `compressed-tensors` package for
its configuration parsing and layer matching functionality.

The following types of quantization are supported in this PR:

- W8A16 and W4A16 INT using GPTQ-Marlin kernels.
- W8A8 and W8A16 FP using FP8-Marlin and cutlass kernels.

Support for other quantization types will be added in subsequent PRs.

a7850008

16 Oct, 2024 1 commit

Fp8 e4m3_fnuz support for rocm (#2588) · 704a58c8

Mohit Sharma authored Oct 16, 2024

* (feat) fp8 fnuz support for rocm

* (review comments) Fix compression_config load, type hints

* (bug) update all has_tensor

* (review_comments) fix typo and added comments

* (nit) improved comment

704a58c8

08 Oct, 2024 1 commit

Add support for fused MoE Marlin for AWQ (#2616) · 64142489

Daniël de Kok authored Oct 08, 2024

* Add support for fused MoE Marlin for AWQ

This uses the updated MoE Marlin kernels from vLLM.

* Add integration test for AWQ MoE

64142489

30 Sep, 2024 1 commit
- MoE Marlin: support `desc_act` for `groupsize != -1` (#2590) · 1c84a30f
  Daniël de Kok authored Sep 30, 2024
```
This change uses the updated Marlin MoE kernel from vLLM to support
MoE with activation sorting and groups.
```
  1c84a30f
31 Jul, 2024 1 commit

Handle GPTQ-Marlin loading in `GPTQMarlinWeightLoader` (#2300) · 34f7dcfd

Daniël de Kok authored Jul 31, 2024

The `GPTWeightLoader` was structured like this in pseudocode:

if marlin:
  Set up tensors in a way that GPTQ-Marlin expects
else:
  Set up tensors in a way that ExLlama/GPTQ/AWQ expect

However, the GPT-Marlin implementation details should really be in the
`marlin` module. So move the former part out to a separate
`GPTQMarlinWeightsLoader`.

34f7dcfd

29 Jul, 2024 1 commit
- Install Marlin from standalone package (#2320) · 922732b2
  Daniël de Kok authored Jul 29, 2024
  
  922732b2
26 Jul, 2024 1 commit

feat: add ruff and resolve issue (#2262) · bab02ff2

drbh authored Jul 26, 2024

* feat: add ruff and resolve issue

* fix: update client exports and adjust after rebase

* fix: adjust syntax to avoid circular import

* fix: adjust client ruff settings

* fix: lint and refactor import check and avoid model enum as global names

* fix: improve fbgemm_gpu check and lints

* fix: update lints

* fix: prefer comparing model enum over str

* fix: adjust lints and ignore specific rules

* fix: avoid unneeded quantize check

bab02ff2

24 Jul, 2024 1 commit
- Split up `layers.marlin` into several files (#2292) · 93d2b9fe
  Daniël de Kok authored Jul 24, 2024
```
The marlin.py file was getting large, split it up.
```
  93d2b9fe