Commits · 92abe6033491dcaa958235e551f40f6b417d3771 · chenpangpang / transformers

31 Jul, 2024 1 commit

>3-5x faster torch.compile forward compilation for autoregressive decoder models (#32227) · 92abe603

fxmarty authored Jul 31, 2024



* draft

* apply changes to all relevant archs

* rerun ci - check_docstrings.py failing?

* fix docstring

* move 2D->4D mask creation to modeling file

* repo consistency

* fix the batch size = 1 case - calling contiguous is not enough

* nit

* style

* propagate to gemma/gemma-2

* prepare inputs for gemma generation

* implement test and tiny fix in gemma2

* Update src/transformers/models/bloom/modeling_bloom.py
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* fix copies

* ci pass

* fix gemma's test_compile_static_cache tests

* flacky

* retrigger ci

---------
Co-authored-by: sanchit-gandhi <sanchit@huggingface.co>
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

92abe603

30 Jul, 2024 1 commit

Docs: formatting nits (#32247) · e68ec18c

Joao Gante authored Jul 30, 2024



* doc formatting nits

* ignore non-autodocs

* Apply suggestions from code review
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Update src/transformers/models/esm/modeling_esm.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Update src/transformers/models/esm/modeling_esm.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* make fixup

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

e68ec18c

26 Jul, 2024 1 commit

Adds: extra_repr for RMSNorm layers in most models (#32204) · f9756d9e

Rohit Dwivedula authored Jul 26, 2024

* adds: extra_repr() to RMSNorm layers in multiple models

* adds: extra_repr for deprecated models as well

* formatting as per style guide

f9756d9e

24 Jul, 2024 1 commit

let's not warn when someone is running a forward (#32176) · 8d2534c4

Arthur authored Jul 24, 2024

* let's not warn when someone is running a foward without cache + self.training

* more models

* fixup

8d2534c4

23 Jul, 2024 2 commits

Enhancing SFT Training Efficiency Using Packing and FlashAttention2 with Position IDs (#31629) · 9cf4f2aa

RhuiDih authored Jul 23, 2024

* add DataCollatorBatchFlattening

* Update data_collator.py

* change name

* new FA2 flow if position_ids is provided

* add comments

* minor fix

* minor fix data collator

* add test cases for models

* add test case for data collator

* remove extra code

* formating for ruff check and check_repo.py

* ruff format

ruff format tests src utils

* custom_init_isort.py

9cf4f2aa

Llama: RoPE refactor (#32135) · 2e113422

Joao Gante authored Jul 23, 2024


Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

2e113422

14 Jul, 2024 1 commit

Generate: remove deprecated code due to `Cache` and `cache_position` being default (#31898) · 739a6316

Joao Gante authored Jul 14, 2024



* tmp commit

* shorter

* nit

* explicit kwargs

* propagate changes

* mass propagation with a few manual touches (let's see how CI behaves)

* fix cacheless case

* Update src/transformers/generation/utils.py
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* make fixup

---------
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

739a6316

11 Jul, 2024 1 commit

Refactor flash attention implementation in transformers (#31446) · e3143952

Arthur authored Jul 11, 2024



* dumb commit

* nit

* update

* something like this

* unpack in modeling utils

* safe import

* oups

* update

* nits

* diff convert gemma

* update

* start propagating

* udpate other modeling code as well

* update for sliding window models

* nits

* more init cleanups

* styling

* fixup

* noice

* pass fixup

* typo typing_extension -> typing_extensions

* torch.nn.functionnal -> torch.nn.functional

* add to import structure

* unpack

* simplify a bit more for this first version

* nut

* update

* update

* nit

* ease the import of `Unpack`

* remove useless `use_sliding_window`

* no qua please

* protect import?

* style

* [run-slow]

* [run slow] llama,gemma,mistral,mixtral

* remove extra kwargs

* fix llama

* address review comments

* apply diff_model_converter to modeling_gemma.py

* remove cache_position 1

* remove cache_position 2

* some cleaning

* refactor gemma2 as well

* apply review comments

* rename file to modeling_flash_attention_utils.py

* siglip refactor

* remove dead code

* is the hub down?

* still down?

* fix siglip

* fix gemma2

* fatal: Could not read from remote repository.

* fix typo in softcap implem

* flacky

* Failed: Timeout >120.0s

---------
Co-authored-by: fxmarty <9808326+fxmarty@users.noreply.github.com>

e3143952

26 Jun, 2024 1 commit

Llama et al. / FSDP : Fix breaking change in 4.40 for FSDP (#31161) · 3f93fd06

Younes Belkada authored Jun 26, 2024



* fix llama fsdp

* fixup

* adding FSDP tests for CPU offloading

* fixes

* fix tests

* fix tests

* add it for mixtral

* propagate the changes on other models

* Update src/transformers/models/phi/modeling_phi.py

* Delete utils/testing_scripts/fsdp_cpu_offloading.py

Remove script - FSDP + CPU offloading it tested in the test suite

* Delete utils/testing_scripts/dummy_fsdp_config.yml

* Update + add cache_positions docstring

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

3f93fd06

21 Jun, 2024 1 commit

Deprecate legacy cache + use cache position (#31491) · 730a4407

Raushan Turganbay authored Jun 21, 2024



* tmp

* update models

* revert utils

* delete

* Update src/transformers/models/dbrx/modeling_dbrx.py
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* modify warning msg

---------
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

730a4407

18 Jun, 2024 1 commit
- Fix typing errors in `Qwen2ForTokenClassification` (#31440) · 1f9387d3
  Kevin Hu authored Jun 18, 2024
```
* Update modeling_qwen2.py

* Fix llama

* More fixes
```
  1f9387d3
05 Jun, 2024 1 commit

Reduce by 2 the memory requirement in `generate()`

🔥

(#30536) · bd5091df

Cyril Vallez authored Jun 05, 2024

* Fix contrastive_search for new cache structure, and improve performance by removing inneficient torch.stack(torch.split(x, top_k, dim=0))

* Fix _contrastive_search for non-standard cache using ellipsis slicing

* Fix all outputs.logits memory leaks for all decoding strategies!

* Fix small error in _contrastive_search()

* Make all necessary change and revert for the new class

* Apply coding style

* Remove pipes in type hints for compatibility

* correct type hint

* apply style

* Use DynamicCache by default and solve conflicts

* Fix rebase issues

* Add `_supports_dynamic_cache_class` in models for models that support DynamicCache but not other caches to make DynamicCache the default for more models

* Create generation config to return legacy format by default, or to choose not to

* style

* Fix case when use_cache is False

* Remove default DynamicCache in assiste_decoding if assistant_model does not support it + fix _seen_tokens when cropping cache

* Update prepare_inputs_for_generation() for case with empty DynamicCache

* Correct return of args in _assisted_decoding

* Remove EfficientDynamicCache as it is no longer needed

* Correct mistake in generation config

* Move cache logic of assisted decoding to AssistedCandidateGenerator.__init__

* change DynamicCache function names from "split" to "batch_split" for readability + apply coding style

* Remove `_supports_dynamic_cache_class` attribute after rebase

* Correct missing line lost in conflict resolution during rebasing

* Add special case for Jamba

* Fix jamba test

* Coding style

* coding style

* Correct missing import in rebasing

* Simplify _validate_model_kwargs based on removal of _supports_dynamic_cache attribute

* Simplify code paths in _contrastive_search

* coding style

* Update docstrings of cache methods

* Update prepare_inputs_for_generation() -> past_key_values are always Cache objects

bd5091df

22 May, 2024 1 commit

update ruff version (#30932) · 673440d0

Arthur authored May 22, 2024



* update ruff version

* fix research projects

* Empty

* Fix errors

---------
Co-authored-by: Lysandre <lysandre@huggingface.co>

673440d0

20 May, 2024 3 commits

Add torch.compile for Mistral (#30642) · 616bb11d

Longjie Zheng authored May 20, 2024

* first version

* fix sliding window

* fix style

* add sliding window cache

* fix style

* address comments

* fix test

* fix style

* move sliding window check inside cache init

* revert changes on irrelevant files & add comment on SlidingWindowCache

* address comments & fix style

fix style

* update causal mask

* [run-slow] mistral

* [run-slow] mistral

* [run-slow] mistral

* [run-slow] mistral

* [run-slow] mistral

* [run-slow] llama

* [run-slow] mistral

* [run-slow] mistral

* [run-slow] mistral

* revert CI from a10 to t4

* wrap up

616bb11d

Add support for torch.compile dynamic shapes (#30560) · cd6bd0af

Benjamin Warner authored May 20, 2024

* add torch.compile dynamic support

* Add SDPA dynamic shapes compile test & improve SDPA comment

* comment consistency

cd6bd0af

Add TokenClassification for Mistral, Mixtral and Qwen2 (#29878) · 07bf2dff

Joseph Enguehard authored May 20, 2024



* Add MistralForTokenClassification

* Add tests and docs

* Add token classification for Mixtral and Qwen2

* Save llma for token classification draft

* Add token classification support for Llama, Gemma, Persimmon, StableLm and StarCoder2

* Formatting

* Add token classification support for Qwen2Moe model

* Add dropout layer to each ForTokenClassification model

* Add copied from in tests

* Update src/transformers/models/llama/modeling_llama.py
Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com>

* Propagate suggested changes

* Style

---------
Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com>

07bf2dff

17 May, 2024 1 commit

Remove deprecated logic and warnings (#30743) · 57c965a8

amyeroberts authored May 17, 2024

* Remove deprecated logic and warnings

* Add back some code that seems to be important...

* Let's just add all he nllb stuff back; removing it is a bit more involved

* Remove kwargs

* Remove more kwargs

57c965a8

16 May, 2024 1 commit
- Cache: add new flag to distinguish models that `Cache` but not static cache (#30800) · 9d889f87
  Joao Gante authored May 16, 2024
```
* jamba cache

* new flag

* generate exception
```
  9d889f87
14 May, 2024 1 commit
- CI: more models wo cache support (#30780) · d8f8a9cd
  Joao Gante authored May 14, 2024
  
  d8f8a9cd
30 Apr, 2024 1 commit
- Cache: Static cache as a standalone object (#30476) · 75bbfd5b
  Joao Gante authored Apr 30, 2024
  
  75bbfd5b
17 Apr, 2024 2 commits

Enable fx tracing for Mistral (#30209) · 304c6a1e
Raushan Turganbay authored Apr 17, 2024
```
* tracing for mistral

* typo

* fix copies
```
304c6a1e

Fix SDPA sliding window compatibility (#30127) · 40eb6d6c

fxmarty authored Apr 17, 2024



* fix sdpa + sliding window

* give credit
Co-authored-by: ehuaa <ehuamail@163.com>

* remove unnecessary warning

* fix typog

* add test

---------
Co-authored-by: ehuaa <ehuamail@163.com>

40eb6d6c

05 Apr, 2024 1 commit
- Fix mixtral ONNX Exporter Issue. (#29858) · d704c0b6
  Adam Louly authored Apr 05, 2024
```
* fix mixtral onnx export

* fix qwen model
```
  d704c0b6
27 Mar, 2024 1 commit

MixtralSparseMoeBlock: add gate jitter (#29865) · a25037be

Lorenzo Verardo authored Mar 27, 2024

This commit adds gate jitter to MixtralSparseMoeBlock's input data
before passing it through the MoE layer, if turned on.

a25037be

08 Mar, 2024 1 commit

StableLM: Fix dropout argument type error (#29236) · f386c51a

liangjs authored Mar 08, 2024



* fix stablelm dropout argument type error

* fix docs of _flash_attention_forward

* fix all docs of _flash_attention_forward

* fix docs of _flash_attention_forward in starcoder2

---------
Co-authored-by: oliang <oliang@tencent.com>

f386c51a

04 Mar, 2024 1 commit
- [Mixtral] Fixes attention masking in the loss (#29363) · 39ef3fb2
  Siming Dai authored Mar 04, 2024
```
Fix mixtral load balancing loss
Co-authored-by: dingkunbo <dingkunbo@baidu.com>
```
  39ef3fb2
28 Feb, 2024 1 commit

Disable Mixtral `output_router_logits` during inference (#29249) · 2ce56d35

Leonardo Emili authored Feb 28, 2024

* Set output_router_logits=False in prepare_inputs_for_generation for mixtral

* Add output_router_logits=False to prepare_inputs_for_generation for mixtral

* Fix style

2ce56d35

14 Feb, 2024 1 commit
- [`CLeanup`] Revert SDPA attention changes that got in the static kv cache PR (#29027) · 609a1767
  Arthur authored Feb 15, 2024
```
* revert unrelated changes that got in

* style
```
  609a1767
08 Feb, 2024 2 commits

fix: torch.int32 instead of torch.torch.int32 (#28883) · 0b693e90
vodkaslime authored Feb 08, 2024

0b693e90

[`Core generation`] Adds support for static KV cache (#27931) · 115ac94d

Arthur authored Feb 08, 2024

Co-authored-by: fxmarty <9808326+fxmarty@users.noreply.github.com>
Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com>
Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>

115ac94d

31 Jan, 2024 1 commit
- DeepSpeed: hardcode `torch.arange` dtype on `float` usage to avoid incorrect... · beb2a096
  Joao Gante authored Jan 31, 2024
```
DeepSpeed: hardcode `torch.arange` dtype on `float` usage to avoid incorrect initialization (#28760)
```
  beb2a096
29 Jan, 2024 1 commit
- Fix typo of `Block`. (#28727) · e694e985
  xkszltl authored Jan 29, 2024
  
  e694e985
24 Jan, 2024 1 commit

Exclude the load balancing loss of padding tokens in Mixtral-8x7B (#28517) · c5c69096

Khai Mai authored Jan 24, 2024

* fix the function load_balancing_loss_func in Mixtral_Moe to include attention_mask

* format code using black and ruff

* skip computing mask if attention_mask=None

* add tests for load balancing loss Mixtral-Moe

* fix assert loss is different in mixtral_test

* fix pad_leng

* use assertNotAlmostEqual and print to debug

* remove print for debug

* minor updates

* reduce rtol and atol

c5c69096

15 Jan, 2024 1 commit
- [`chore`] Update warning text, a word was missing (#28017) · 881e966a
  Tom Aarsen authored Jan 15, 2024
```
Update warning, a word was missing
```
  881e966a
12 Jan, 2024 1 commit
- Docs: add model paths (#28475) · dc01cf9c
  Joao Gante authored Jan 12, 2024
  
  dc01cf9c
11 Jan, 2024 1 commit

Fix load balancing loss func for mixtral (#28256) · e768616a

liangxuZhang authored Jan 11, 2024



* Correct the implementation of auxiliary loss of mixtrtal

* correct the implementation of auxiliary loss of mixtrtal

* Implement a simpler calculation method

---------
Co-authored-by: zhangliangxu3 <zhangliangxu3@jd.com>

e768616a

05 Jan, 2024 1 commit
- chore: Fix typo s/exclusivelly/exclusively/ (#28361) · 4ab5fb89
  hugo-syn authored Jan 05, 2024
  
  4ab5fb89
26 Dec, 2023 1 commit
- fix FA2 when using quantization (#28203) · 3b7675b2
  Sourab Mangrulkar authored Dec 26, 2023
  
  3b7675b2
22 Dec, 2023 1 commit

Fix ONNX export for causal LM sequence classifiers by removing reverse indexing (#28144) · 548a8f61

Dean Wyatte authored Dec 22, 2023

* normalize reverse indexing for causal lm sequence classifiers

* normalize reverse indexing for causal lm sequence classifiers

* normalize reverse indexing for causal lm sequence classifiers

* use modulo instead

* unify modulo-based sequence lengths

548a8f61

21 Dec, 2023 1 commit

[`Mixtral` & `Mistral`] Add support for sdpa (#28133) · f9a98c47

Arthur authored Dec 21, 2023



* some nits

* update test

* add support d\sd[a

* remove some dummy inputs

* all good

* style

* nits

* fixes

* fix more copies

* nits

* styling

* fix

* Update src/transformers/models/mistral/modeling_mistral.py
Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com>

* add a slow test just to be sure

* fixup

---------
Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com>

f9a98c47