Commits · b026b5ca6dd178f50d91778f58d32d85d3a83c5f · chenpangpang / transformers

06 Nov, 2023 1 commit
- Fix tokenizer export for LLamaTokenizerFast (#27222) · b026b5ca
  Mayank Mishra authored Nov 06, 2023
```
* fix tokenizer

* fix tokenizer
```
  b026b5ca
03 Nov, 2023 12 commits

translate run_scripts.md to chinese (#27246) · cc3e4781

jiaqiw09 authored Nov 03, 2023

* translate run_scripts.md to chinese

* translate run_scripts.md to chinese

* translate run_scripts.md to chinese

cc3e4781

translate autoclass_tutorial to chinese (#27269) · bf7cfac2
jiaqiw09 authored Nov 03, 2023
```
* translate autoclass_tutorial.md  to chinese

* translate update
```
bf7cfac2

[`FA2`] Add flash attention for for `DistilBert` (#26489) · 1ac2463d

Susnato Dhar authored Nov 03, 2023

* flash attention added for DistilBert

* fixes

* removed padding_masks

* Update modeling_distilbert.py

* Update test_modeling_distilbert.py

* style fix

1ac2463d

[Docs] Model_doc structure/clarity improvements (#26876) · 5964f820

Maria Khalusova authored Nov 03, 2023

* first batch of structure improvements for model_docs

* second batch of structure improvements for model_docs

* more structure improvements for model_docs

* more structure improvements for model_docs

* structure improvements for cv model_docs

* more structural refactoring

* addressed feedback about image processors

5964f820

[`Docs` / `SAM` ] Reflect correct changes to run inference without OOM (#27268) · ad8ff962
Younes Belkada authored Nov 03, 2023
```
Update sam.md
```
ad8ff962
Fix switch transformer mixed precision issue (#27220) · f13f544a
Shiyu Li authored Nov 03, 2023
```
* Fix mixed precision error for switch transformer

* Fixup
```
f13f544a

Update the ConversationalPipeline docstring for chat templates (#27250) · db69bd88

Matt authored Nov 03, 2023

* Update the ConversationalPipeline docstring now that we're using chat templates

* Direct access to conversation.messages

* Explain the string init

db69bd88

[docs] Custom model doc update (#27213) · 011b15c1
Maria Khalusova authored Nov 03, 2023
```
doc update
```
011b15c1

Avoid many failing tests in doctesting (#27262) · af8d1dc3

Yih-Dar authored Nov 03, 2023



* fix

* update

* update

* fix

---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

af8d1dc3

[`PEFT` / `Tests` ] Fix peft integration failing tests (#27258) · 8f1a43cd
Younes Belkada authored Nov 03, 2023
```
fix peft integration issues
```
8f1a43cd

Refactor: Use Llama RoPE implementation for Falcon (#26933) · 05ea7b79

Tom Aarsen authored Nov 03, 2023

* Use Llama RoPE implementation for Falcon

+ Add copy functionalities

* Use standard cache format for Falcon

* Simplify apply_rotary_pos_emb, copy from Llama

* Remove unnecessary cache conversion test

We don't need to convert any caches anymore!

* Resolve copy complaint

05ea7b79

Fuyu protection (#27248) · e9a6c72b
Lysandre Debut authored Nov 03, 2023

e9a6c72b

02 Nov, 2023 15 commits

Fixed base model class name extraction from PeftModels (#27162) · 552ff244

Komal Kumar authored Nov 02, 2023

* Fixed base model class name extraction from PeftModels

* Changes to first unwrap the model then extract the base model name

* Changed base_model to base_model.model to stay consistent with peft model abstractions

552ff244

Removed the redundant SiLUActivation class. (#27136) · 49912168

Chi authored Nov 02, 2023

* Removed the redundant SiLUActivation class and now use nn.functional.silu directly.

* I apologize for adding torch.functional.silu. I have replaced it with nn.SiLU.

49912168

translate peft.md to chinese (#27215) · 00d8502b
jiaqiw09 authored Nov 02, 2023
```
* tranlsate peft.md to chinese

* translate peft.md to chinese

* fix missing link
```
00d8502b
Dev version · bc78fd12
Lysandre authored Nov 02, 2023

bc78fd12

Enrich TTS pipeline parameters naming (#26473) · 0ed6729b

Yoach Lacombe authored Nov 02, 2023



* enrich TTS pipeline docstring for clearer forward_params use

* change token leghts

* update Pipeline parameters

* correct docstring and make style

* fix tests

* make style

* change music prompt
Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com>

* Apply suggestions from code review
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>
Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com>

* raise errors if generate_kwargs with forward-only models

* make style

---------
Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com>
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

0ed6729b

Remove redundant code from T5 encoder mask creation (#27216) · 147e8ce4
Pietro Lesci authored Nov 02, 2023
```
* remove redundant code

* update

* add typecasting

* make `attention_mask` float again
```
147e8ce4
Generate: return `past_key_values` (#25086) · a6c82d45
Joao Gante authored Nov 02, 2023

a6c82d45
fix-deprecated-exllama-arg (#27243) · 441c3e0d
Marc Sun authored Nov 02, 2023
```
fix-exllama
```
441c3e0d

Fixing m4t. (#27240) · 8801861d

Nicolas Patry authored Nov 02, 2023

* Fixing m4t.

* Trying to remove comparison ? Odd test failure.

* Adding shared. But why on earth does it hang ????

* Putting back the model weights checks the test is silently failing on
cuda.

* Fix style + unremoved comment.

8801861d

Fix safetensors failing tests (#27231) · 443bf5e9

Lysandre Debut authored Nov 02, 2023



* Fix Kosmos2

* Fix ProphetNet

* Fix MarianMT

* Fix M4T

* XLM ProphetNet

* ProphetNet fix

* XLM ProphetNet

* Final M4T fixes

* Tied weights keys

* Revert M4T changes

* Apply suggestions from code review
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

443bf5e9

Wrap `_prepare_4d_causal_attention_mask` as a leaf function (#27236) · 4557a0de
Michael Benayoun authored Nov 02, 2023
```
Wrap _prepare_4d_causal_attention_mask as a leaf function
```
4557a0de

Fuyu: improve image processing (#27007) · 8a312956

Pablo Montalvo authored Nov 02, 2023



* Fix Fuyu image scaling bug

It could produce negative padding and hence inference errors for certain
image sizes.

* initial rework commit

* add batching capabilities, refactor image processing

* add functional batching for a list of images and texts

* make args explicit

* Fuyu processing update (#27133)

* Add file headers

* Add file headers

* First pass - preprocess method with standard args

* First pass image processor rework

* Small tweaks

* More args and docstrings

* Tidying iterating over batch

* Tidying up

* Modify to have quick tests (for now)

* Fix up

* BatchFeature

* Passing tests

* Add tests for processor

* Sense check when patchifying

* Add some tests

* FuyuBatchFeature

* Post-process box coordinates

* Update to `size` in processor

* Remove unused and duplicate constants

* Store unpadded dims after resize

* Fix up

* Return FuyuBatchFeature

* Get unpadded sizes after resize

* Update exception

* Fix return

* Convert input `<box>` coordinates to model format.

* Post-process point coords, support multiple boxes/points in a single
sequence

* Replace constants

* Update src/transformers/models/fuyu/image_processing_fuyu.py
Co-authored-by: Pedro Cuenca <pedro@huggingface.co>

* Preprocess List[List[image]]

* Update src/transformers/models/fuyu/image_processing_fuyu.py
Co-authored-by: Pedro Cuenca <pedro@huggingface.co>

* Update to Amy's latest state.

* post-processing returns a list of tensors

* Fix error when target_sizes is None
Co-authored-by: Pablo Montalvo <pablo.montalvo.leroux@gmail.com>

* Update src/transformers/models/fuyu/image_processing_fuyu.py
Co-authored-by: Pedro Cuenca <pedro@huggingface.co>

* Update src/transformers/models/fuyu/image_processing_fuyu.py
Co-authored-by: Pedro Cuenca <pedro@huggingface.co>

* Update src/transformers/models/fuyu/image_processing_fuyu.py
Co-authored-by: Pedro Cuenca <pedro@huggingface.co>

* Update src/transformers/models/fuyu/image_processing_fuyu.py
Co-authored-by: Pedro Cuenca <pedro@huggingface.co>

* Review comments

* Update src/transformers/models/fuyu/image_processing_fuyu.py
Co-authored-by: Pedro Cuenca <pedro@huggingface.co>

* Fix up

* Fix up

---------
Co-authored-by: Ubuntu <ubuntu@ip-172-31-72-126.ec2.internal>
Co-authored-by: Pedro Cuenca <pedro@huggingface.co>
Co-authored-by: Pablo Montalvo <pablo.montalvo.leroux@gmail.com>

* Fix conflicts in fuyu_follow_up_image_processing (#27228)

fixing conflicts and updating on main

* Revert "Fix conflicts in fuyu_follow_up_image_processing" (#27232)

Revert "Fix conflicts in fuyu_follow_up_image_processing (#27228)"

This reverts commit acce10b6c653dc7041fb9d18cfed55775afd6207.

---------
Co-authored-by: Pedro Cuenca <pedro@huggingface.co>
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>
Co-authored-by: Ubuntu <ubuntu@ip-172-31-72-126.ec2.internal>

8a312956

[`core` / `Quantization`] Fix for 8bit serialization tests (#27234) · 9b25c164
Younes Belkada authored Nov 02, 2023
```
* fix for 8bit serialization

* added regression tests.

* fixup
```
9b25c164

Reproducible checkpoint for npu (#27208) · c52e429b

Hz, Ji authored Nov 02, 2023

* save NPU's RNG states when saving a checkpoint and set after all the
data skip phase when resuming training.

* re-trigger ci

* re-trigger ci

c52e429b

support bf16 (#25879) · 7adaefe2

Roohollah Etemadi authored Nov 02, 2023

* added bf16 support

* added cuda availability check

* applied make style, quality

7adaefe2

01 Nov, 2023 12 commits

[Whisper, Bart, MBart] Add Flash Attention 2 (#27203) · af3de8d8

Patrick von Platen authored Nov 01, 2023



* add whisper fa2

* correct

* change all

* correct

* correct

* fix more

* fix more

* fix more

* fix more

* fix more

* fix more

* Apply suggestions from code review
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* fix more

* fix more

* fix more

* fix more

* fix more

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

af3de8d8

Enable split_batches through TrainingArguments (#26798) · 3520e37e

Zach Mueller authored Nov 01, 2023

* Enable split_batches through TrainingArguments

* Extra dispatch_batches

* Keep as default false

* Add to docstring

* Add to docstring

* Remove the capturewarnings change

* Comma

3520e37e

Fix CPU offload + disk offload tests (#27204) · 95020f20
Lysandre Debut authored Nov 01, 2023
```
Fix disk offload tests + weight sharing issues
```
95020f20

Add exllamav2 better (#27111) · c9e72f55

Marc Sun authored Nov 01, 2023



* add_ xllamav2 arg

* add test

* style

* add check

* add doc

* replace by use_exllama_v2

* fix tests

* fix doc

* style

* better condition

* fix logic

* add deprecate msg

* deprecate exllama

* remove disable_exllama from the linter

* remove

* fix warning

* Revert the commits deprecating exllama

* deprecate disable_exllama for use_exllama

* fix

* fix loading attribute

* better handling of args

* remove disable_exllama from init and linter

* Apply suggestions from code review
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* better arg

* fix warning

* Apply suggestions from code review
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* switch to dict

* Apply suggestions from code review
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* style

* nits

* style

* better tests

* style

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

c9e72f55

Translate task summary to chinese (#27180) · 239cd0ea

jiaqiw09 authored Nov 01, 2023

* translate task_summary.md to chinese

* update translation

* update translation

* fix _toctree.yml

239cd0ea

improving TimmBackbone to support FrozenBatchNorm2d (#27160) · 1e32b05e

Rafael Padilla authored Nov 01, 2023



* supporting freeze_batch_norm_2d

* supporting freeze_batch_norm_2d

* including unfreeze + separate into methods

* fix typo

* calling unfreeze

* lint

* Update src/transformers/models/timm_backbone/modeling_timm_backbone.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

---------
Co-authored-by: Rafael Padilla <rafael.padilla@huggingface.co>
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

1e32b05e

Fix docstring in get_oneformer_resize_output_image_size func (#27207) · 21a2fbaf
Wesley L Passos authored Nov 01, 2023

21a2fbaf

Add TensorFlow implementation of ConvNeXTv2 (#25558) · f8afb2b2

Andi Powers Holmes authored Nov 02, 2023

* Add type annotations to TFConvNextDropPath

* Use tf.debugging.assert_equal for TFConvNextEmbeddings shape check

* Add TensorFlow implementation of ConvNeXTV2

* check_docstrings: add TFConvNextV2Model to exclusions

TFConvNextV2Model and TFConvNextV2ForImageClassification have docstrings
which are equivalent to their PyTorch cousins, but a parsing issue prevents them
from passing the test.

Adding exclusions for these two classes as discussed in #25558.

f8afb2b2

[WhisperForCausalLM] Add WhisperForCausalLM for speculative decoding (#27195) · 391d14e8

Patrick von Platen authored Nov 01, 2023



* finish

* add tests

* fix all tests

* [Assistant Decoding] Add test

* fix more

* better

* finish

* Apply suggestions from code review
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* finish

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

391d14e8

Added cache_block_outputs option to enable GPTQ for non-regular models (#27032) · f9b4bea0

Alexander Kozlov authored Nov 01, 2023



* Added cache_block_outputs option to enable GPTQ for non-regular models

* Update src/transformers/utils/quantization_config.py
Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>

* Update src/transformers/utils/quantization_config.py
Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>

* Fixed style

* Update src/transformers/utils/quantization_config.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

---------
Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

f9b4bea0

added unsqueeze_dim to apply_rotary_pos_emb (#27117) · 037fb7d0

Shashank Rajput authored Nov 01, 2023



* added unsqueeze_dim to apply_rotary_pos_emb

* Added docstring

* Modified docstring

* Modified docstring

* Modified docstring

* Modified docstring

* Modified docstring

* ran make fix-copies and make fixup

* Update src/transformers/models/llama/modeling_llama.py

Accepting the proposed changes in formatting.
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* incorporating PR suggestions

* incorporating PR suggestions

* incorporating PR suggestions

* incorporating PR suggestions

* ..

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

037fb7d0

Fixing docstring in get_resize_output_image_size function (#27191) · f3c1a172
Wesley L Passos authored Nov 01, 2023

f3c1a172