Commits · 38b53da38af231b0af967d15ca29c52470e402d5 · chenpangpang / transformers

26 Apr, 2024 8 commits

[examples] update whisper fine-tuning (#29938) · 38b53da3

Sanchit Gandhi authored Apr 26, 2024

* [examples] update whisper fine-tuning

* deprecate forced/suppress tokens

* item assignment

* update readme

* final fix

38b53da3

[`DETR`] Remove timm hardcoded logic in modeling files (#29038) · aafa7ce7

amyeroberts authored Apr 26, 2024



* Enable instantiating model with pretrained backbone weights

* Clarify pretrained import

* Use load_backbone instead

* Add backbone_kwargs to config

* Fix up

* Add tests

* Tidy up

* Enable instantiating model with pretrained backbone weights

* Update tests so backbone checkpoint isn't passed in

* Clarify pretrained import

* Update configs - docs and validation check

* Update src/transformers/utils/backbone_utils.py
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* Clarify exception message

* Update config init in tests

* Add test for when use_timm_backbone=True

* Use load_backbone instead

* Add use_timm_backbone to the model configs

* Add backbone_kwargs to config

* Pass kwargs to constructors

* Draft

* Fix tests

* Add back timm - weight naming

* More tidying up

* Whoops

* Tidy up

* Handle when kwargs are none

* Update tests

* Revert test changes

* Deformable detr test - don't use default

* Don't mutate; correct model attributes

* Add some clarifying comments

* nit - grammar is hard

---------
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

aafa7ce7

Remove skipping logic now that set_epoch exists (#30501) · 77ff304d
Zach Mueller authored Apr 26, 2024
```
* Remove skipping logic now that set_epoch exists

* Working version, clean
```
77ff304d

[`BERT`] Add support for sdpa (#28802) · dfa7b580

JB (Don) authored Apr 26, 2024

* Adding SDPA support for BERT

* Using the proper input name for testing model input in inference()

* Adding documentation for SDPA in BERT model page

* Use the stable link for the documentation

* Adding a gate to only call .contiguous() for torch < 2.2.0

* Additions and fixes to the documentation

* Minor updates to documentation

* Adding extra requirements needed for the contiguous() bug

* Adding "Adapted from" in plcae of the "Copied from"

* Add benchmark speedup tables to the documentation

* Minor fixes to the documentation

* Use ClapText as a replacemenet for Bert in the Copied-From

* Some more fixes for the fix-copies references

* Overriding the test_eager_matches_sdpa_generate in bert tests to not load with low_cpu_mem_usage

[test all]

* Undo changes to separate test

* Refactored SDPA self attention code for KV projections

* Change use_sdpa to attn_implementation

* Fix test_sdpa_can_dispatch_on_flash by preparing input (required for MultipleChoice models)

dfa7b580

Use the Keras set_random_seed in tests (#30504) · 2de5cb12
Matt authored Apr 26, 2024
```
Use the Keras set_random_seed to ensure reproducible weight initialization
```
2de5cb12
Update `dtype_byte_size` to handle torch.float8_e4m3fn/float8_e5m2 types (#30488) · 20081c74
Michael Goin authored Apr 26, 2024
```
* Update modeling_utils/dtype_byte_size to handle float8 types

* Add a test for dtype_byte_size

* Format

* Fix bool
```
20081c74
Fix the `bitsandbytes` error formatting ("Some modules are dispatched on ...") (#30494) · 59e715f7
kyo authored Apr 26, 2024
```
Fix the `bitsandbytes` error when some modules are not properly offloaded.
```
59e715f7
FEAT: PEFT support for EETQ (#30449) · 19cfdf0f
Younes Belkada authored Apr 26, 2024
```
Update quantizer_eetq.py
```
19cfdf0f

25 Apr, 2024 18 commits

[docs] Spanish translation of pipeline_tutorial.md (#30252) · a98c4179

Aaron Jimenez authored Apr 25, 2024

* add pipeline_webserver to es/

* add pipeline_webserver to es/, translate first section

* add comment for checking link

* translate pipeline_webserver

* edit pipeline_webserver

* fix typo

a98c4179

Quantization: `HfQuantizer` quant method update (#30484) · 26ddc580
Younes Belkada authored Apr 25, 2024
```
ensure popular quant methods are supported
```
26ddc580

Add sidebar tutorial for chat models (#30401) · f3962712

Matt authored Apr 25, 2024



* Draft tutorial for talking to chat models

* Reformat lists and text snippets

* Cleanups and clarifications

* Finish up remaining TODOs

* Correct section link

* Small fix

* Add proper quantization examples

* Add proper quantization examples

* Add proper quantization examples

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/_toctree.yml
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Update docs/source/en/conversations.md
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Fix Text Generation Pipeline link and add a ref to the LLM inference guide

* intelligent -> capable

* Small intro cleanup

* Small text cleanup

* Small text cleanup

* Clarification about system message

* Clarification about system message

---------
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

f3962712

Do not use deprecated `SourceFileLoader.load_module()` in dynamic module loading (#30370) · bc274a28
Xuehai Pan authored Apr 26, 2024

bc274a28
Fix Llava for 0-embeddings (#30473) · e60491ad
Raushan Turganbay authored Apr 25, 2024

e60491ad

Introduce Stateful Callbacks (#29666) · ad697f18

Zach Mueller authored Apr 25, 2024



* Introduce saveable callbacks

* Add note

* Test for non-present and flag

* Support early stopping and refusing to train further

* Update docstring

* More saving

* Import oopsie

* Apply suggestions from code review
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Make it go through TrainerArguments

* Document

* Fix test

* Apply suggestions from code review
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Rework to allow for duplicates

* CLean

* Fix failing tests

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

ad697f18

Make accelerate install non-torch dependent (#30463) · 86f25697

Zach Mueller authored Apr 25, 2024



* Pin accelerate w/o eager

* Eager

* Update .circleci/create_circleci_config.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Expound

* Expound squared

* PyTorch -> dependency

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

86f25697

Fix Issue #29817 Video Classification Task Guide Using Undeclared Variables (#30457) · 92833138

manju rangam authored Apr 25, 2024



* Fix issue #29817

Video Classification Task Guide Using Undeclared Variables

* Update docs/source/en/tasks/video_classification.md

updated with review comments
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Fix issue #29817

Add line space following PR comments

---------
Co-authored-by: manju-rangam <Manju1@Git>
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

92833138

Add WSD scheduler (#30231) · 7b1170b0

Alexander Visheratin authored Apr 25, 2024

* Added WSD scheduler.

* Added tests.

* Fixed errors.

* Fix formatting.

* CI fixes.

7b1170b0

🚨

Add training compatibility for Musicgen-like models (#29802) · 90cb55bf

Yoach Lacombe authored Apr 25, 2024



* first modeling code

* make repository

* still WIP

* update model

* add tests

* add latest change

* clean docstrings and copied from

* update docstrings md and readme

* correct chroma function

* correct copied from and remove unreleated test

* add doc to toctree

* correct imports

* add convert script to notdoctested

* Add suggestion from Sanchit
Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com>

* correct get_uncoditional_inputs docstrings

* modify README according to SANCHIT feedback

* add chroma to audio utils

* clean librosa and torchaudio hard dependencies

* fix FE

* refactor audio decoder -> audio encoder for consistency with previous musicgen

* refactor conditional -> encoder

* modify sampling rate logics

* modify license at the beginning

* refactor all_self_attns->all_attentions

* remove ignore copy from causallm generate

* add copied from for from_sub_models

* fix make copies

* add warning if audio is truncated

* add copied from where relevant

* remove artefact

* fix convert script

* fix torchaudio and FE

* modify chroma method according to feedback-> better naming

* refactor input_values->input_features

* refactor input_values->input_features and fix import fe

* add input_features to docstrigs

* correct inputs_embeds logics

* remove dtype conversion

* refactor _prepare_conditional_hidden_states_kwargs_for_generation ->_prepare_encoder_hidden_states_kwargs_for_generation

* change warning for chroma length

* Update src/transformers/models/musicgen_melody/convert_musicgen_melody_transformers.py
Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com>

* change way to save wav, using soundfile

* correct docs and change to soundfile

* fix import

* fix init proj layers

* add draft training

* fix cross entropy

* clean loss computation

* fix labels

* remove line breaks from md

* fix issue with docstrings

* add FE suggestions

* improve is in logics and remove useless imports

* remove custom from_pretrained

* simplify docstring code

* add suggestions for modeling tests

* make style

* update converting script with sanity check

* remove encoder attention mask from conditional generation

* replace musicgen melody checkpoints with official orga

* rename ylacombe->facebook in checkpoints

* fix copies

* remove unecessary warning

* add shape in code docstrings

* add files to slow doc tests

* fix md bug and add md to not_tested

* make fix-copies

* fix hidden states test and batching

* update training code

* add training tests for melody

* add training for o.g musicgen

* fix copied from

* remove final todos

* make style

* fix style

* add suggestions from review

* add ref to the original loss computation code

* rename method + fix labels in tests

* make style

---------
Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com>

90cb55bf

Prevent crash with `WandbCallback` with third parties (#30477) · ce5ae5a4

Tom Aarsen authored Apr 25, 2024

* Use EAFP principle to prevent crash with third parties

* Remove leftover debugging code

* Add info-level logger message

ce5ae5a4

Don't run fp16 MusicGen tests on CPU (#30466) · aca4a103
amyeroberts authored Apr 25, 2024

aca4a103

Fix SigLip classification doctest (#30475) · 4fed29e3

amyeroberts authored Apr 25, 2024

* Fix SigLip classification doctest

* Remove extra line

* Update src/transformers/models/siglip/modeling_siglip.py

4fed29e3

Script for finding candidate models for deprecation (#29686) · 30ee508c

amyeroberts authored Apr 25, 2024

* Add utility for finding candidate models for deprecation

* Better model filtering

* Update

* Add warning tip

* Fix up

* Review comments

* Filter requests based on tags

* Add copyright header

30ee508c

[fix codellama conversion] (#30472) · c60749d6
Arthur authored Apr 25, 2024
```
* fix codellama conversion

* nit
```
c60749d6
FIX / Workflow: Fix SSH workflow bug (#30474) · e9b16354
Younes Belkada authored Apr 25, 2024
```
Update ssh-runner.yml
```
e9b16354
FIX / Workflow: Change tailscale trigger condition (#30471) · cd0cd12a
Younes Belkada authored Apr 25, 2024
```
Update push-important-models.yml
```
cd0cd12a

Workflow / ENH: Add SSH into our runners workflow (#30425) · cebb0726

Younes Belkada authored Apr 25, 2024



* add SSH into our runners workflow

* fix

* fix

* fix

* use our previous approaches

* forward contrib credits from discussions

---------
Co-authored-by: Yih-Dar <ydshieh@users.noreply.github.com>

cebb0726

24 Apr, 2024 14 commits

consistent job / pytest report / artifact name correspondence (#30392) · fbb41cd4

Yih-Dar authored Apr 24, 2024



* better names

* run better names

* update

* update

---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

fbb41cd4

Non blocking support to torch DL's (#30465) · 6ad9c8f7
Zach Mueller authored Apr 24, 2024
```
* Non blocking support

* Check for optimization

* Doc
```
6ad9c8f7

Enable fp16 on CPU (#30459) · 5c57463b

Zach Mueller authored Apr 24, 2024

* Check removing flag for torch

* LLM oops

* Getting there...

* More discoveries

* Change

* Clean up and prettify

* Logic check

* Not

5c57463b

Neuron: When save_safetensor=False, no need to move model to CPU (#29703) · d1d94d79

jeffhataws authored Apr 24, 2024

save_safetensor=True is default as of release 4.35.0, which then
required TPU hotfix https://github.com/huggingface/transformers/pull/27799
(issue https://github.com/huggingface/transformers/issues/27578).
However, when the flag save_safetensor is set to False (compatibility mode),
moving the model to CPU causes generation of too many graphs
during checkpoint https://github.com/huggingface/transformers/issues/28438.
This PR disable moving of model to CPU when save_safetensor=False.

d1d94d79

[`research_project`] Most of the security issues come from this requirement.txt (#29977) · 661190b4
Arthur authored Apr 24, 2024
```
update most of decision transformers research project
```
661190b4
Fix wrong indent in `utils/check_if_new_model_added.py` (#30456) · d0d430f1
Yih-Dar authored Apr 24, 2024
```
fix
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
d0d430f1

Phi-3 (#30423) · c9693db2

Gustavo de Rosa authored Apr 24, 2024

* chore(root): Initial commit of Phi-3 files.

* fix(root): Fixes Phi-3 missing on readme.

* fix(root): Ensures files are consistent.

* fix(phi3): Fixes unit tests.

* fix(tests): Fixes style of phi-3 test file.

* chore(tests): Adds integration tests for Phi-3.

* fix(phi3): Removes additional flash-attention usage, .e.g, swiglu and rmsnorm.

* fix(phi3): Fixes incorrect docstrings.

* fix(phi3): Fixes docstring typos.

* fix(phi3): Adds support for Su and Yarn embeddings.

* fix(phi3): Improves according first batch of reviews.

* fix(phi3): Uses up_states instead of y in Phi3MLP.

* fix(phi3): Uses gemma rotary embedding to support torch.compile.

* fix(phi3): Improves how rotary embedding classes are defined.

* fix(phi3): Fixes inv_freq not being re-computed for extended RoPE.

* fix(phi3): Adds last suggestions to modeling file.

* fix(phi3): Splits inv_freq calculation in two lines.

c9693db2

Add `paths` filter to avoid the chance of being triggered (#30453) · 42fed15c
Yih-Dar authored Apr 24, 2024
```
* trigger

* remove the last job

---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
42fed15c

[SegGPT] Fix loss calculation (#30421) · d26c1413

Eduardo Pacheco authored Apr 24, 2024



* Fixed main train issues

* Added loss test

* Update src/transformers/models/seggpt/modeling_seggpt.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Added missing labels arg in SegGptModel forward

* Fixed typo

* Added slow test to test loss calculation

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

d26c1413

fix jamba slow foward for multi-gpu (#30418) · 37fa1f65
Marc Sun authored Apr 24, 2024
```
* fix jamba slow foward for multi-gpu

* remove comm

* oups

* style
```
37fa1f65
fix uncaught init of linear layer in clip's/siglip's for image classification models (#30435) · 5d64ae9d
Anton Vlasjuk authored Apr 24, 2024
```
* fix clip's/siglip's _init_weights to reflect linear layers in "for image classification"

* trigger slow tests
```
5d64ae9d
[tests] make test device-agnostic (#30444) · 16c8e176
Fanli Lin authored Apr 24, 2024
```
* make device-agnostic

* clean code
```
16c8e176

[`Llava`] + CIs fix red cis and llava integration tests (#30440) · 9a4a119c

Arthur authored Apr 24, 2024



* nit

* nit and fmt skip

* fixup

* Update src/transformers/convert_slow_tokenizer.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* set to true

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

9a4a119c

Fix YOLOS image processor resizing (#30436) · 767e3518

Pavel Iakubovskii authored Apr 24, 2024

* Add test for square image that fails

* Fix for square images

* Extend test cases

* Fix resizing in tests

* Style fixup

767e3518