Commits · 0ae789e04330e15a90e34cd723c851a8ab8d7ec5 · chenpangpang / transformers

30 Apr, 2024 2 commits

Enable multi-device for more models (#30409) · 0ae789e0

Jacky Lee authored Apr 30, 2024

* feat: support for dinov2

* feat: support for depth_anything

* feat: support for efficientformer

* feat: support for bert (is this right?)

* update: embedding split

* remove: empty string

* feat: support for align

* fix: copies

* fix: QAQBertEmbeddings

* fix: more consistency issues

* revert: support for effientformer

* feat: support for altclip

* feat: support for blip_text

* support for ChineseCLIP

* feat: support for depth anything

* feat: support for dpt

* feat: support for dpt

* feat: support for git

* feat: support for groupvit

* update: format

* fix: support for clip

* fix: consistency

* feat: support for pvt

* feat: support for vit_msn

* fix: consistency

* fix: other copies

* remove: device transfer

* revert: in-place add

* update: support for align

* update: support for bert

* update: support for Chinese CLIP

* revert: changes to efficientformer

* update: support for dpt

* update: support for efficientformer

* revert: changes to git

* revert: changes to groupvit

* revert: changes to roc_bert

* update: support for vit_msn

* revert: changes to dpt

* remove: extra space

* style: extra space

0ae789e0

Pass `use_cache` in kwargs for GPTNeoX (#30538) · c712d05a
Raushan Turganbay authored Apr 30, 2024
```
pass use_cache in kwargs
```
c712d05a

29 Apr, 2024 7 commits
- Include safetensors as part of `_load_best_model` (#30553) · a3aabc70
  Zach Mueller authored Apr 29, 2024
```
* Include safetensors

* Cleanup
```
  a3aabc70
- Reenable SDPA's FA2 During Training with torch.compile (#30442) · 9df8b301
  Benjamin Warner authored Apr 29, 2024
```
* Reenable SDPA's FA2 during training with torch.compile

* fix Olmo's SDPA FA2 dispatching too

* update formatting

* improved SDPA comment

* formatting and explanatory comment

* is_causal if statement to one-liner
```
  9df8b301
- Fix repo. fetch/checkout in PR slow CI job (#30537) · 87be06ca
  Yih-Dar authored Apr 29, 2024
```
fix
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
  87be06ca
- Update runner tag for PR slow CI (#30535) · c0242188
  Yih-Dar authored Apr 29, 2024
```
fix
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
  c0242188
- Fix broken link to Transformers notebooks (#30512) · bdbe1662
  clinty authored Apr 29, 2024
```
Co-authored-by: Clint Adams <clint@debian.org>
```
  bdbe1662
- Pass attn_implementation when using AutoXXX.from_config (#30507) · e8acb700
  amyeroberts authored Apr 29, 2024
```
* Pass attn_implementation when using AutoXXX.from_config

* Fix
```
  e8acb700
- Allow boolean FSDP options in fsdp_config (#30439) · 80126f98
  Howard Liberty authored Apr 29, 2024
```
* Allow boolean FSDP options in fsdp_config

* Use lower() to be safe
```
  80126f98
26 Apr, 2024 12 commits

Fix link in dbrx.md (#30509) · 73014b56
Eitan Turok authored Apr 26, 2024

73014b56

[SegGPT] Fix seggpt image processor (#29550) · 6d4cabda

Eduardo Pacheco authored Apr 26, 2024

* Fixed SegGptImageProcessor to handle 2D and 3D prompt mask inputs

* Added new test to check prompt mask equivalence

* New proposal

* Better proposal

* Removed unnecessary method

* Updated seggpt docs

* Introduced do_convert_rgb

* nits

6d4cabda

load_image - decode b64encode and encodebytes strings (#30192) · c793b26f
amyeroberts authored Apr 26, 2024
```
* Decode b64encode and encodebytes strings

* Remove conditional encode -- image is always a string
```
c793b26f
Fix GroundingDINO, DPR after BERT SDPA update (#30506) · e7d52a10
amyeroberts authored Apr 26, 2024
```
Fix GroundingDINO, DPR after BET SDPA update
```
e7d52a10

[examples] update whisper fine-tuning (#29938) · 38b53da3

Sanchit Gandhi authored Apr 26, 2024

* [examples] update whisper fine-tuning

* deprecate forced/suppress tokens

* item assignment

* update readme

* final fix

38b53da3

[`DETR`] Remove timm hardcoded logic in modeling files (#29038) · aafa7ce7

amyeroberts authored Apr 26, 2024



* Enable instantiating model with pretrained backbone weights

* Clarify pretrained import

* Use load_backbone instead

* Add backbone_kwargs to config

* Fix up

* Add tests

* Tidy up

* Enable instantiating model with pretrained backbone weights

* Update tests so backbone checkpoint isn't passed in

* Clarify pretrained import

* Update configs - docs and validation check

* Update src/transformers/utils/backbone_utils.py
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* Clarify exception message

* Update config init in tests

* Add test for when use_timm_backbone=True

* Use load_backbone instead

* Add use_timm_backbone to the model configs

* Add backbone_kwargs to config

* Pass kwargs to constructors

* Draft

* Fix tests

* Add back timm - weight naming

* More tidying up

* Whoops

* Tidy up

* Handle when kwargs are none

* Update tests

* Revert test changes

* Deformable detr test - don't use default

* Don't mutate; correct model attributes

* Add some clarifying comments

* nit - grammar is hard

---------
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

aafa7ce7

Remove skipping logic now that set_epoch exists (#30501) · 77ff304d
Zach Mueller authored Apr 26, 2024
```
* Remove skipping logic now that set_epoch exists

* Working version, clean
```
77ff304d

[`BERT`] Add support for sdpa (#28802) · dfa7b580

JB (Don) authored Apr 26, 2024

* Adding SDPA support for BERT

* Using the proper input name for testing model input in inference()

* Adding documentation for SDPA in BERT model page

* Use the stable link for the documentation

* Adding a gate to only call .contiguous() for torch < 2.2.0

* Additions and fixes to the documentation

* Minor updates to documentation

* Adding extra requirements needed for the contiguous() bug

* Adding "Adapted from" in plcae of the "Copied from"

* Add benchmark speedup tables to the documentation

* Minor fixes to the documentation

* Use ClapText as a replacemenet for Bert in the Copied-From

* Some more fixes for the fix-copies references

* Overriding the test_eager_matches_sdpa_generate in bert tests to not load with low_cpu_mem_usage

[test all]

* Undo changes to separate test

* Refactored SDPA self attention code for KV projections

* Change use_sdpa to attn_implementation

* Fix test_sdpa_can_dispatch_on_flash by preparing input (required for MultipleChoice models)

dfa7b580

Use the Keras set_random_seed in tests (#30504) · 2de5cb12
Matt authored Apr 26, 2024
```
Use the Keras set_random_seed to ensure reproducible weight initialization
```
2de5cb12
Update `dtype_byte_size` to handle torch.float8_e4m3fn/float8_e5m2 types (#30488) · 20081c74
Michael Goin authored Apr 26, 2024
```
* Update modeling_utils/dtype_byte_size to handle float8 types

* Add a test for dtype_byte_size

* Format

* Fix bool
```
20081c74
Fix the `bitsandbytes` error formatting ("Some modules are dispatched on ...") (#30494) · 59e715f7
kyo authored Apr 26, 2024
```
Fix the `bitsandbytes` error when some modules are not properly offloaded.
```
59e715f7
FEAT: PEFT support for EETQ (#30449) · 19cfdf0f
Younes Belkada authored Apr 26, 2024
```
Update quantizer_eetq.py
```
19cfdf0f

25 Apr, 2024 18 commits

[docs] Spanish translation of pipeline_tutorial.md (#30252) · a98c4179

Aaron Jimenez authored Apr 25, 2024

* add pipeline_webserver to es/

* add pipeline_webserver to es/, translate first section

* add comment for checking link

* translate pipeline_webserver

* edit pipeline_webserver

* fix typo

a98c4179

Quantization: `HfQuantizer` quant method update (#30484) · 26ddc580
Younes Belkada authored Apr 25, 2024
```
ensure popular quant methods are supported
```
26ddc580

Add sidebar tutorial for chat models (#30401) · f3962712

Matt authored Apr 25, 2024



* Draft tutorial for talking to chat models

* Reformat lists and text snippets

* Cleanups and clarifications

* Finish up remaining TODOs

* Correct section link

* Small fix

* Add proper quantization examples

* Add proper quantization examples

* Add proper quantization examples

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/_toctree.yml
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Fix Text Generation Pipeline link and add a ref to the LLM inference guide

* intelligent -> capable

* Small intro cleanup

* Small text cleanup

* Small text cleanup

* Clarification about system message

* Clarification about system message

---------
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

f3962712

Do not use deprecated `SourceFileLoader.load_module()` in dynamic module loading (#30370) · bc274a28
Xuehai Pan authored Apr 26, 2024

bc274a28
Fix Llava for 0-embeddings (#30473) · e60491ad
Raushan Turganbay authored Apr 25, 2024

e60491ad

Introduce Stateful Callbacks (#29666) · ad697f18

Zach Mueller authored Apr 25, 2024



* Introduce saveable callbacks

* Add note

* Test for non-present and flag

* Support early stopping and refusing to train further

* Update docstring

* More saving

* Import oopsie

* Apply suggestions from code review
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Make it go through TrainerArguments

* Document

* Fix test

* Apply suggestions from code review
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Rework to allow for duplicates

* CLean

* Fix failing tests

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

ad697f18

Make accelerate install non-torch dependent (#30463) · 86f25697

Zach Mueller authored Apr 25, 2024



* Pin accelerate w/o eager

* Eager

* Update .circleci/create_circleci_config.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Expound

* Expound squared

* PyTorch -> dependency

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

86f25697

Fix Issue #29817 Video Classification Task Guide Using Undeclared Variables (#30457) · 92833138

manju rangam authored Apr 25, 2024



* Fix issue #29817

Video Classification Task Guide Using Undeclared Variables

* Update docs/source/en/tasks/video_classification.md

updated with review comments
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Fix issue #29817

Add line space following PR comments

---------
Co-authored-by: manju-rangam <Manju1@Git>
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

92833138

Add WSD scheduler (#30231) · 7b1170b0

Alexander Visheratin authored Apr 25, 2024

* Added WSD scheduler.

* Added tests.

* Fixed errors.

* Fix formatting.

* CI fixes.

7b1170b0

🚨

Add training compatibility for Musicgen-like models (#29802) · 90cb55bf

Yoach Lacombe authored Apr 25, 2024



* first modeling code

* make repository

* still WIP

* update model

* add tests

* add latest change

* clean docstrings and copied from

* update docstrings md and readme

* correct chroma function

* correct copied from and remove unreleated test

* add doc to toctree

* correct imports

* add convert script to notdoctested

* Add suggestion from Sanchit
Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com>

* correct get_uncoditional_inputs docstrings

* modify README according to SANCHIT feedback

* add chroma to audio utils

* clean librosa and torchaudio hard dependencies

* fix FE

* refactor audio decoder -> audio encoder for consistency with previous musicgen

* refactor conditional -> encoder

* modify sampling rate logics

* modify license at the beginning

* refactor all_self_attns->all_attentions

* remove ignore copy from causallm generate

* add copied from for from_sub_models

* fix make copies

* add warning if audio is truncated

* add copied from where relevant

* remove artefact

* fix convert script

* fix torchaudio and FE

* modify chroma method according to feedback-> better naming

* refactor input_values->input_features

* refactor input_values->input_features and fix import fe

* add input_features to docstrigs

* correct inputs_embeds logics

* remove dtype conversion

* refactor _prepare_conditional_hidden_states_kwargs_for_generation ->_prepare_encoder_hidden_states_kwargs_for_generation

* change warning for chroma length

* Update src/transformers/models/musicgen_melody/convert_musicgen_melody_transformers.py
Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com>

* change way to save wav, using soundfile

* correct docs and change to soundfile

* fix import

* fix init proj layers

* add draft training

* fix cross entropy

* clean loss computation

* fix labels

* remove line breaks from md

* fix issue with docstrings

* add FE suggestions

* improve is in logics and remove useless imports

* remove custom from_pretrained

* simplify docstring code

* add suggestions for modeling tests

* make style

* update converting script with sanity check

* remove encoder attention mask from conditional generation

* replace musicgen melody checkpoints with official orga

* rename ylacombe->facebook in checkpoints

* fix copies

* remove unecessary warning

* add shape in code docstrings

* add files to slow doc tests

* fix md bug and add md to not_tested

* make fix-copies

* fix hidden states test and batching

* update training code

* add training tests for melody

* add training for o.g musicgen

* fix copied from

* remove final todos

* make style

* fix style

* add suggestions from review

* add ref to the original loss computation code

* rename method + fix labels in tests

* make style

---------
Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com>

90cb55bf

Prevent crash with `WandbCallback` with third parties (#30477) · ce5ae5a4

Tom Aarsen authored Apr 25, 2024

* Use EAFP principle to prevent crash with third parties

* Remove leftover debugging code

* Add info-level logger message

ce5ae5a4

Don't run fp16 MusicGen tests on CPU (#30466) · aca4a103
amyeroberts authored Apr 25, 2024

aca4a103

Fix SigLip classification doctest (#30475) · 4fed29e3

amyeroberts authored Apr 25, 2024

* Fix SigLip classification doctest

* Remove extra line

* Update src/transformers/models/siglip/modeling_siglip.py

4fed29e3

Script for finding candidate models for deprecation (#29686) · 30ee508c

amyeroberts authored Apr 25, 2024

* Add utility for finding candidate models for deprecation

* Better model filtering

* Update

* Add warning tip

* Fix up

* Review comments

* Filter requests based on tags

* Add copyright header

30ee508c

[fix codellama conversion] (#30472) · c60749d6
Arthur authored Apr 25, 2024
```
* fix codellama conversion

* nit
```
c60749d6
FIX / Workflow: Fix SSH workflow bug (#30474) · e9b16354
Younes Belkada authored Apr 25, 2024
```
Update ssh-runner.yml
```
e9b16354
FIX / Workflow: Change tailscale trigger condition (#30471) · cd0cd12a
Younes Belkada authored Apr 25, 2024
```
Update push-important-models.yml
```
cd0cd12a

Workflow / ENH: Add SSH into our runners workflow (#30425) · cebb0726

Younes Belkada authored Apr 25, 2024



* add SSH into our runners workflow

* fix

* fix

* fix

* use our previous approaches

* forward contrib credits from discussions

---------
Co-authored-by: Yih-Dar <ydshieh@users.noreply.github.com>

cebb0726

24 Apr, 2024 1 commit

consistent job / pytest report / artifact name correspondence (#30392) · fbb41cd4

Yih-Dar authored Apr 24, 2024



* better names

* run better names

* update

* update

---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

fbb41cd4