Commits · f0735f95174136a71a097ce54942c1e9a9d89a3a · OpenDAS / bitsandbytes

05 Dec, 2024 1 commit

LLM.int8() Refactoring: Part 1 (#1401) · 81e6345d

Matthew Douglas authored Dec 05, 2024



* Start of int8 refactor: remove col32/col_ampere/col_turing transforms in new igemmlt implementation

* Fix unintended change

* New naive mm_dequant kernel for row-major; cleanup

* fix

* int8 refactor: initial sparse decomp, cleanup

* Int8 refactoring: remove separate NO_CUBLASLT build; more cleanup

* int8: inference optimizations, some cleanup

* int8: more tests passing, cleanup

* int8 - more cleanup, most tests passing

* int8: specify CUDA stream for int8 ops

* perf: reduce overhead from getting cudaStream ptr

* Mark some functions for deprecation.

* int8 sparse decomp: small perf improvement

* update setup.py

* Update bitsandbytes/autograd/_functions.py
Co-authored-by: Aarni Koskela <akx@iki.fi>

* Update bitsandbytes/functional.py
Co-authored-by: Aarni Koskela <akx@iki.fi>

* Update bitsandbytes/functional.py
Co-authored-by: Aarni Koskela <akx@iki.fi>

* Update bitsandbytes/research/autograd/_functions.py
Co-authored-by: Aarni Koskela <akx@iki.fi>

* int8 - perf improvement for sparse decomposition inference; deprecate get_tensor_stream() in favor of new private fn

* int8 cleanup

* Ignore ruff rule ISC001 (incompatible with formatter)

* add comment

* int8 more cleanup

* Update bitsandbytes/functional.py
Co-authored-by: Aarni Koskela <akx@iki.fi>

* int8: rename / deprecate old fn signatures

* Update bitsandbytes/functional.py
Co-authored-by: Aarni Koskela <akx@iki.fi>

* type annotation

* format update

* Update bitsandbytes/research/autograd/_functions.py
Co-authored-by: Aarni Koskela <akx@iki.fi>

* cleanup

* Add comment to explain division optimization

* more cleanup

* Update bitsandbytes/functional.py
Co-authored-by: Aarni Koskela <akx@iki.fi>

* Update bitsandbytes/functional.py
Co-authored-by: Aarni Koskela <akx@iki.fi>

* Update bitsandbytes/functional.py
Co-authored-by: Aarni Koskela <akx@iki.fi>

* cleanup

* Type annotations, cleanup

* remove unused kernels; improved type annotations

* small perf optimization for single-GPU systems

* small perf optimization for single-GPU systems

* update docstrings

* Improve docs and tests

* Update docstring

* Update test

* add benchmarking script

* test cleanup: add deprecated marker, move benchmarks out

* Add int8 dequant function; misc improvements

* int8 matmul fallback for inner dims not divisible by 4

* improve register usage of kInt8VectorQuant - especially for A100/H100

* disable fail-fast for package build

* maxwell compat

* ptxas verbose

* docs update

* doc update

* backward fix

* Bugfix sparse decomp

* Int8 fix for PEFT OLoRA init

* Fix test for deprecated spmm_coo

* test improvement

* doc update

* typo

* doc cleanup

* docs

* add inference benchmark script

* Add benchmarks, doc update

---------
Co-authored-by: Aarni Koskela <akx@iki.fi>

81e6345d

29 Mar, 2024 1 commit
- Fix 4bit quantization with blocksize=4096 · c17fb8eb
  Matthew Douglas authored Mar 29, 2024
  
  c17fb8eb
13 Mar, 2024 1 commit
- Reformat with ruff-format · 5a4263f4
  Ruff authored Feb 24, 2024
  
  5a4263f4
21 Feb, 2024 1 commit
- tests: fix all_close to respect max 2 positional args (#1074) · d11b5068
  Titus authored Feb 21, 2024
  
  d11b5068
01 Feb, 2024 3 commits

Enable line-ending and other hygiene lints (#1006) · 6974920b
Aarni Koskela authored Feb 01, 2024

6974920b

Test improvements (#1001) · 2336a45c

Aarni Koskela authored Feb 01, 2024

* test_nvidia_transform: fix variable reference

`out_order` is the global parametrization list, not the test fixture argument

* Make `parametrize` use more idiomatic

* Use a more deterministic helper for `dim*` determination

* Convert NO_CUBLASLT errors into skips too

* Mark slow and benchmark tests as such (allows `-k "not benchmark"`)

2336a45c

test_nvidia_transform: fix variable reference (#1000) · 1a0dc5c3
Aarni Koskela authored Feb 01, 2024
```
`out_order` is the global parametrization list, not the test fixture argument
```
1a0dc5c3

30 Jan, 2024 1 commit

Ruff fixes (#984) · 706ec24d

Aarni Koskela authored Jan 30, 2024



* Adjust Ruff configuration

* do not autofix always
* be less strict around tests and benchmarks
* adjust ignores for now

* Ruff: autofix I and F401

* Apply ruff autofixes

* Fix RUF013 complaint

* Fix mutable default in replace_linear

* Don't use bare except

* Wrap bitsandbytes.__main__ entrypoint in function; fix "sensible" typo

* Fix ruff B008 (function call in arguments)

* Add ruff noqas as suitable

* Fix RUF005 (splat instead of concatenating)

* Fix B018 (useless expression)

* Add pre-commit configuration + GitHub Actions lint workflow

* Fix unused `e` in bitsandbytes/__main__.py

* fix merge conflict resolution error

* run pre-commit hook

---------
Co-authored-by: Titus <9048635+Titus-von-Koeller@users.noreply.github.com>

706ec24d

24 Jan, 2024 1 commit

Tests: improve CUDA support detection (#985) · f1c75741

Aarni Koskela authored Jan 24, 2024

* implicitly skip any test that implicitly uses CUDA on a non-CUDA box
* add a `requires_cuda` fixture

f1c75741

17 Jan, 2024 1 commit

Initial FSDP Support for QLoRA Finetuning (#970) · dcfb6f81

Benjamin Warner authored Jan 16, 2024



This PR adds initial FSDP support for training QLoRA models. It enables basic FSDP and CPU Offload support, with low memory training via FSDP.sync_module_states option unsupported.

This PR builds off of #840 commit 8278fca and BNB FSDP by @TimDettmers and @Titus-von-Koeller.

An example of using this PR to finetune QLoRA models with FSDP can be found in the demo repo: AnswerDotAi/fsdp_qlora.

* Minimal changes for fp32 4bit storage from BNB commit 8278fca

* Params4bit with selectable storage dtype

* possible fix for double quantizing linear weight & quant storage dtype

* minor fixes in Params4bit for peft tests

* remove redundant

* add float16

* update test

* Remove float16 quant cast as there are fp32, bf16, & fp16 quant kernels

---------
Co-authored-by: Kerem Turgutlu <keremturgutlu@gmail.com>

dcfb6f81

08 Jan, 2024 1 commit
- Fixed bnb input in setup.py. Bumped version for release. · 4870580f
  Tim Dettmers authored Jan 07, 2024
  
  4870580f
02 Nov, 2023 2 commits
- reverted fn signatures in functional() · 4c11d6dc
  Ruslan Svirschevski authored Sep 20, 2023
  
  4c11d6dc
- use QuantState class for quant_state · 61a4a20d
  Ruslan Svirschevski authored Sep 11, 2023
  
  61a4a20d
04 Aug, 2023 1 commit
- Fixed two bugs in dynamic data type creation. · 3c9aca91
  Tim Dettmers authored Aug 03, 2023
  
  3c9aca91
19 Jul, 2023 1 commit
- Increased occupancy. · c82f51c0
  Tim Dettmers authored Jul 19, 2023
  
  c82f51c0
12 Jul, 2023 1 commit
- Fixed missing bias in bnb.matmul_4bit for inference; more tests. · 90b0ac57
  Tim Dettmers authored Jul 11, 2023
  
  90b0ac57
11 Jul, 2023 1 commit
- Added more extensive gemv tests; blocksize guard for gemv. · ba51d95d
  Tim Dettmers authored Jul 11, 2023
  
  ba51d95d
10 Jul, 2023 3 commits
- Removed debugging statement. · a26a321e
  Tim Dettmers authored Jul 10, 2023
  
  a26a321e
- Fixed accidential deletion of limits in kernel. · 306f6b23
  Tim Dettmers authored Jul 10, 2023
  
  306f6b23
- Added fp32 compute type for gemv_4bit. · 5fab6734
  Tim Dettmers authored Jul 09, 2023
  
  5fab6734
09 Jul, 2023 3 commits
- Added double quantization support and tests. · 0f0390ac
  Tim Dettmers authored Jul 09, 2023
  
  0f0390ac
- Added FP4 fast inference support. · 94168d79
  Tim Dettmers authored Jul 09, 2023
  
  94168d79
- Added abitrary data types; fixed a bug for small matrices. · 4b88d69d
  Tim Dettmers authored Jul 09, 2023
  
  4b88d69d
08 Jul, 2023 2 commits
- Turning optimization (float accumulation). 185 vs 50. · eefbf602
  Tim Dettmers authored Jul 08, 2023
  
  eefbf602
- Added warp_shuffle indexing 185 vs 54. · 7e49b5b9
  Tim Dettmers authored Jul 08, 2023
  
  7e49b5b9
05 Jul, 2023 1 commit
- Added bfloat16 quantizations and tests. · 02fd80cb
  Tim Dettmers authored Jul 04, 2023
  
  02fd80cb
04 Jul, 2023 2 commits
- Vectorized loads, conflict free NF4; 52 vs 172. · dfe6900b
  Tim Dettmers authored Jul 04, 2023
  
  dfe6900b
- Initial 4-bit naive batch size 1, 81 vs 185. · f89ff93e
  Tim Dettmers authored Jul 03, 2023
  
  f89ff93e
31 May, 2023 2 commits
- Added debugging functions. · e54d2730
  Tim Dettmers authored May 30, 2023
  
  e54d2730
- Added lookup table. · b7f04e2a
  Tim Dettmers authored May 30, 2023
  
  b7f04e2a
24 May, 2023 1 commit
- Fixed Makefile. · 2bce175d
  Tim Dettmers authored May 23, 2023
  
  2bce175d
06 May, 2023 2 commits
- Added paged optimizers. · 44d68ff2
  Tim Dettmers authored May 06, 2023
  
  44d68ff2
- Added paging. · ec38ba95
  Tim Dettmers authored May 06, 2023
  
  ec38ba95
02 May, 2023 7 commits
- 4-bit draft; 128 vector load 240. · 264a9485
  Tim Dettmers authored May 02, 2023
  
  264a9485
- Warp multi-specialization 240. · 869b7e83
  Tim Dettmers authored May 02, 2023
  
  869b7e83
- Shared memory efficient 240. · 77f15fdc
  Tim Dettmers authored May 02, 2023
  
  77f15fdc
- Correct implementation 240. · 394749db
  Tim Dettmers authored May 02, 2023
  
  394749db
- Initial. · 9aa232cc
  Tim Dettmers authored May 02, 2023
  
  9aa232cc
- Tighter and scaled error analysis. · 9192c9de
  Tim Dettmers authored May 02, 2023
  
  9192c9de
- Baseline for debugging. · f9bfea8f
  Tim Dettmers authored May 02, 2023
  
  f9bfea8f