Commits · b8b16475d41b66ab0e1fe9d1cb82bbff65e5f6d6 · chenpangpang / transformers

"vscode:/vscode.git/clone" did not exist on "fa876aee2adf525b597495c10ad9c96896953dbd"

20 Feb, 2024 7 commits

Revert low cpu mem tie weights (#29135) · 0996a100

amyeroberts authored Feb 20, 2024

* Revert "Add tie_weights() to LM heads and set bias in set_output_embeddings() (#28948)"

This reverts commit 725f4ad1.

* Revert "Patch to skip failing `test_save_load_low_cpu_mem_usage` tests (#29043)"

This reverts commit 4156f517.

0996a100

[`Core tokenization`] `add_dummy_prefix_space` option to help with latest issues (#28010) · 15cfe389

Arthur authored Feb 20, 2024

* add add_dummy_prefix_space option to slow

* checking kwargs might be better. Should be there for all spm tokenizer IMO

* nits

* fix copies

* more copied

* nits

* add prefix space

* nit

* nits

* Update src/transformers/convert_slow_tokenizer.py

* fix inti

* revert wrong styling

* fix

* nits

* style

* updates

* make sure we use slow tokenizer for conversion instead of looking for the decoder

* support llama ast well

* update llama tokenizer fast

* nits

* nits nits nits

* update the doc

* update

* update to fix tests

* skip unrelated tailing test

* Update src/transformers/convert_slow_tokenizer.py

* add proper testing

* test decode as well

* more testing

* format

* fix llama test

* Apply suggestions from code review

15cfe389

FIX [`PEFT` / `Trainer` ] Handle better peft + quantized compiled models (#29055) · efdd4366
Younes Belkada authored Feb 20, 2024
```
* handle peft + compiled models

* add tests

* fixup

* adapt from suggestions

* clarify comment
```
efdd4366
Generate: unset GenerationConfig parameters do not raise warning (#29119) · a7755d24
Joao Gante authored Feb 20, 2024

a7755d24
Llama: fix batched generation (#29109) · 7d312ad2
Joao Gante authored Feb 20, 2024

7d312ad2
FIX [`bnb` / `tests`] Propagate the changes from #29092 to 4-bit tests (#29122) · ff76e7c2
Younes Belkada authored Feb 20, 2024
```
* forgot to push the changes for 4bit ..

* trigger CI
```
ff76e7c2

FEAT [`Trainer` / `bnb`]: Add RMSProp from `bitsandbytes` to HF `Trainer` (#29082) · f7ef7cec

Younes Belkada authored Feb 20, 2024



* add RMSProp to Trainer

* revert some change

* Update src/transformers/trainer.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

f7ef7cec

19 Feb, 2024 4 commits

Bnb test fix for different hardwares (#29066) · 5ce90f32

Titus authored Feb 19, 2024



* generated text on A10G

* generated text in CI

* Apply suggestions from code review

add explanatory comments
Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com>

---------
Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com>

5ce90f32

ENH: added new output_logits option to generate function (#28667) · 08cd694e

Max Baak authored Feb 19, 2024

output_logits option behaves like output_scores, but returns the raw, unprocessed prediction logit scores,
ie. the values before they undergo logit processing and/or warping. The latter happens by default for the
regular output scores.

It's useful to have the unprocessed logit scores in certain circumstances. For example, unprocessed logit scores
are very useful with causallm models when one wants to determine the probability of a certain answer, e.g.
when asking a question with a yes/no answer. In that case getting the next-token probabilities of both "yes" and
"no" (and/or their relative ratio) is of interest for classification. The reason for getting these _before_ logit
processing and/or warping is b/c a) that can change the probabilities or b) reject the tokens of interest / reduce
the number of tokens to just 1.

For an example use-case see paper TabLLM: Few-shot Classification of Tabular Data with Large Language Models
by Stefan Hegselmann, Alejandro Buendia, Hunter Lang, Monica Agrawal, Xiaoyi Jiang, and David Sontag.
https://arxiv.org/abs/2210.10723



In addition:
- added dedicated unit test: tests/generation/test_utils/test_return_unprocessed_logit_scores
  which tests return of logics with output_logits=True in generation.
- set output_logits=True in all other generation unit tests, that also have output_scores=True.

Implemented @gante's and @amyeroberts review feedback
Co-authored-by: kx79wq <max.baak@ing.com>

08cd694e

Fix the `bert-base-cased` tokenizer configuration test (#29105) · 98308586
Lysandre Debut authored Feb 19, 2024
```
Fix test
```
98308586
FIX [`bnb` / `tests`]: Fix currently failing bnb tests (#29092) · a75a6c93
Younes Belkada authored Feb 19, 2024
```
Update test_mixed_int8.py
```
a75a6c93

16 Feb, 2024 7 commits

Add chat support to text generation pipeline (#28945) · 2f1003be

Matt authored Feb 16, 2024

* Add chat support to text generation pipeline

* Better handling of single elements

* Deprecate ConversationalPipeline

* stash commit

* Add missing add_special_tokens kwarg

* Update chat templating docs to refer to TextGenerationPipeline instead of ConversationalPipeline

* Add ✨TF✨

 tests

* @require_tf

* Add type hint

* Add specific deprecation version

* Remove unnecessary do_sample

* Remove todo - the discrepancy has been resolved

* Update src/transformers/tokenization_utils_base.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Update src/transformers/pipelines/text_generation.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

2f1003be

Fix trainer test wrt DeepSpeed + auto_find_bs (#29061) · 636b0324

Zach Mueller authored Feb 16, 2024



* FIx trainer test

* Update tests/trainer/test_trainer.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

636b0324

Honor trust_remote_code for custom tokenizers (#28854) · be42c24d

Richard Lee authored Feb 16, 2024



* pass through trust_remote_code for dynamically loading unregistered tokenizers specified by config
add test

* change directories back to previous directory after test

* fix ruff check

* Add a note to that block for future in case we want to remove it later

---------
Co-authored-by: Matt <rocketknight1@gmail.com>

be42c24d

fix failing trainer ds tests (#29057) · b2628086
Sourab Mangrulkar authored Feb 16, 2024

b2628086

fix num_assistant_tokens with heuristic schedule (#28759) · 258da40e

Jonathan Mamou authored Feb 16, 2024



* fix heuristic num_assistant_tokens_schedule

* Update src/transformers/generation/configuration_utils.py
Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>

* Update src/transformers/generation/candidate_generator.py
Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>

* Update utils.py

check that candidate_generator.assistant_model exists since some some speculations (like ngram and PLD) don't have assistant_model attribute

* Update src/transformers/generation/candidate_generator.py
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* Update tests/generation/test_utils.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* make fixup

* merge conflict

* fix docstring

* make fixup

---------
Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

258da40e

Fix max_length criteria when using inputs_embeds (#28994) · aee11fe4

Raushan Turganbay authored Feb 16, 2024



* fix max_length for inputs_embeds

* make style

* Update src/transformers/generation/utils.py
Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>

* Static Cache: load models with MQA or GQA (#28975)

* fix

* fix tests

* fix tests

* Update src/transformers/generation/utils.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* more fixes

* make style

---------
Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

aee11fe4

Update all references to canonical models (#29001) · f497f564
Lysandre Debut authored Feb 16, 2024
```
* Script & Manual edition

* Update
```
f497f564

15 Feb, 2024 3 commits

Patch to skip failing `test_save_load_low_cpu_mem_usage` tests (#29043) · 4156f517
amyeroberts authored Feb 15, 2024
```
* Patch to skip currently failing tests

* Whoops - wrong place
```
4156f517

DeformableDetrModel support fp16 (#29013) · 5b6fa230

Donggeun Yu authored Feb 15, 2024



* Update ms_deform_attn_cuda.cu

* Update ms_deform_attn_cuda.cuh

* Update modeling_deformable_detr.py

* Update src/transformers/models/deformable_detr/modeling_deformable_detr.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Update modeling_deformable_detr.py

* python utils/check_copies.py --fix_and_overwrite

* Fix dtype missmatch error

* Update test_modeling_deformable_detr.py

* Update test_modeling_deformable_detr.py

* Update modeling_deformable_detr.py

* Update modeling_deformable_detr.py

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

5b6fa230

Fix static generation when compiling! (#28937) · f3788b09

Arthur authored Feb 15, 2024



* wow I was scared!

* fix everything

* nits

* make it BC?

* add todo

* nits

* is_tracing should still be used to pass tracing tests

* nits

* some nits to make sure genration works with static cache uncompiled

* fix sdpa

* fix FA2 for both static and dynamic in a better way?

* style

* fix-copies

* fix fix copies

* fix sequential beam searcg

* style

* use `keys_to_ignore`

* nit

* correct dtype inference when init

* :( the fix for FA2 is still not optimal to investigate!

* styling

* nits

* nit

* this might work better

* add comment

* Update src/transformers/models/llama/modeling_llama.py

* "position_ids" -> "cache_position"

* style

* nit

* Remove changes that should no be propagatted just yet

* Apply suggestions from code review

* Styling

* make sure we raise an errir for static cache with FA2 enabled

* move  to the bottom of the signature

* style

* Update src/transformers/models/llama/modeling_llama.py
Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com>

* Update src/transformers/models/llama/modeling_llama.py

* nit in the name

---------
Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com>

f3788b09

14 Feb, 2024 9 commits

FIX [`Trainer` / tags]: Fix trainer + tags when users do not pass `"tags"` to... · 7a0fccc6

Younes Belkada authored Feb 14, 2024

FIX [`Trainer` / tags]: Fix trainer + tags when users do not pass `"tags"` to `trainer.push_to_hub()` (#29009)

* fix trainer tags

* add test

7a0fccc6

Backbone kwargs in config (#28784) · 0199a484

amyeroberts authored Feb 14, 2024



* Enable instantiating model with pretrained backbone weights

* Clarify pretrained import

* Use load_backbone instead

* Add backbone_kwargs to config

* Pass kwargs to constructors

* Fix up

* Input verification

* Add tests

* Tidy up

* Update tests/utils/test_backbone_utils.py
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

---------
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

0199a484

Add tie_weights() to LM heads and set bias in set_output_embeddings() (#28948) · 725f4ad1

JB (Don) authored Feb 15, 2024

* Add tie_weights() to LM heads and set bias in set_output_embeddings()

The bias were not tied correctly in some LM heads, and this change should fix that.

* Moving test_save_and_load_low_cpu_mem_usage to ModelTesterMixin

* Adding _tie_weights() to MPNet and Vilt

* Skip test for low cpu mem usage for Deta/DeformableDetr since they cannot init on meta device

* Rename to test name to save_load to match the convention

725f4ad1

Fix flaky test vision encoder-decoder generate (#28923) · 354775bc
Raushan Turganbay authored Feb 14, 2024

354775bc

Introduce AcceleratorConfig dataclass (#28664) · 0507e69d

Zach Mueller authored Feb 14, 2024



* Introduce acceleratorconfig dataclass

* Extra second warn

* Move import

* Try moving import under is_accelerate_available

* Quality

* Apply suggestions from code review
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Clean

* Remove to_kwargs

* Change version

* Improve tests by including dispatch and split batches

* Improve reliability

* Update tests/trainer/test_trainer.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Fixup tests and review nits

* Make tests pass

* protect import

* Protect import

* Empty-Commit

* Make training_args.to_dict handle the AcceleratorConfig

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

0507e69d

Set the dataset format used by `test_trainer` to float32 (#28920) · 69ca640d
Huazhong Ji authored Feb 14, 2024
```
Co-authored-by: unit_test <test@unit.com>
```
69ca640d

AQLM quantizer support (#28928) · 1ecf5f7c

Andrei Panferov authored Feb 14, 2024



* aqlm init

* calibration and dtypes

* docs

* Readme update

* is_aqlm_available

* Simpler link in docs

* Test TODO real reference

* init _import_structure fix

* AqlmConfig autodoc

* integration aqlm

* integrations in tests

* docstring fix

* legacy typing

* Less typings

* More kernels information

* Performance -> Accuracy

* correct tests

* remoced multi-gpu test

* Update docs/source/en/quantization.md
Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com>

* Update src/transformers/utils/quantization_config.py
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* Brought back multi-gpu tests

* Update src/transformers/integrations/aqlm.py
Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>

* Update tests/quantization/aqlm_integration/test_aqlm.py
Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>

---------
Co-authored-by: Andrei Panferov <blacksamorez@yandex-team.ru>
Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com>
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>
Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>

1ecf5f7c

Add SiglipForImageClassification and CLIPForImageClassification (#28952) · 63ffd56d
NielsRogge authored Feb 14, 2024
```
* First draft

* Add CLIPForImageClassification

* Remove scripts

* Fix doctests
```
63ffd56d

Add `StableLM` (#28810) · de6029a0

Jonathan Tow authored Feb 14, 2024

* Add `StableLM`

* fix(model): re-create from `huggingface-cli add-new-model-like persimmon`

* fix: re-add changes to address comments

* fix(readme): add links to paper

* fix(tokenization_auto): remove `GPTNeoXTokenizerFastFast` ref

* fix(tests): re-add `@slow` decorator to integration tests

* fix(tests): import slow...

* fix(readme_hd): remove whitespace edit

* fix(tokenizer): auto tokenizer tuple

* skip doctests for `modeling_stablelm`

de6029a0

13 Feb, 2024 4 commits

[`DETR`] Update the processing to adapt masks & bboxes to reflect padding (#28363) · bd4b83e1

amyeroberts authored Feb 13, 2024

* Update the processing so bbox coords are adjusted for padding

* Just pad masks

* Tidy up, add tests

* Better tests

* Fix yolos and mark as slow for pycocotols

* Fix yolos - return_tensors

* Clarify padding and normalization behaviour

bd4b83e1

Static Cache: load models with MQA or GQA (#28975) · 3e70a207
Joao Gante authored Feb 13, 2024

3e70a207

Add sudachi_projection option to BertJapaneseTokenizer (#28503) · da20209d

Hiroshi Matsuda authored Feb 13, 2024



* add sudachi_projection option

* Upgrade sudachipy>=0.6.8

* add a test case for sudachi_projection

* Compatible with older versions of SudachiPy

* make fixup

* make style

* error message for unidic download

* revert jumanpp test cases

* format options for sudachi_projection
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* format options for sudachi_split_mode and sudachi_dict_type

* comment

* add tests for full_tokenizer kwargs

* pass projection arg directly

* require_sudachi_projection

* make style

* revert upgrade sudachipy

* check is_sudachi_projection_available()

* revert dependency_version_table and bugfix

* style format

* simply raise ImportError
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* simply raise ImportError

---------
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

da20209d

[`NllbTokenizer`] refactor with added tokens decoder (#27717) · b4456753

Arthur authored Feb 13, 2024



* refactor with addedtokens decoder

* style

* get rid of lang code to id

* style

* keep some things for BC

* update tests

* add the mask token at the end of the vocab

* nits

* nits

* fix final tests

* style

* nits

* Update src/transformers/models/nllb/tokenization_nllb_fast.py
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* nits

* style?

* Update src/transformers/convert_slow_tokenizer.py

* make it a tad bit more custom

* ruff please stop
Co-Authored by avidale

<dale.david@mail.ru>

* Update
Co-authored-by: avidale <dale.david@mail.ru>

* Update
Co-authored-by: avidale <dale.david@mail.ru>

* oupts

* ouft

* nites

* test

* fix the remaining failing tests

* style

* fix failing test

* ficx other test

* temp dir + test the raw init

* update test

* style

---------
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

b4456753

12 Feb, 2024 3 commits
- [Docs] Add language identifiers to fenced code blocks (#28955) · fe3df9d5
  Klaus Hipp authored Feb 12, 2024
```
Add language identifiers to code blocks
```
  fe3df9d5
- Tests: tag `test_save_load_fast_init_from_base` as flaky (#28930) · e30bbb26
  Joao Gante authored Feb 12, 2024
  
  e30bbb26
- [Nougat] Fix pipeline (#28242) · f278ef20
  NielsRogge authored Feb 12, 2024
```
* Fix pipeline

* Remove print statements

* Address comments

* Address issue

* Remove unused imports
```
  f278ef20
08 Feb, 2024 2 commits

Support batched input for decoder start ids (#28887) · d6286646

Raushan Turganbay authored Feb 08, 2024



* support batched input for decoder start ids

* Fix typos
Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>

* minor changes

* fix: decoder_start_id as list

* empty commit

* empty commit

* empty commit

* empty commit

* empty commit

* empty commit

* empty commit

* empty commit

* empty commit

---------
Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>

d6286646

[`Core generation`] Adds support for static KV cache (#27931) · 115ac94d

Arthur authored Feb 08, 2024

Co-authored-by: fxmarty <9808326+fxmarty@users.noreply.github.com>
Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com>
Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>

115ac94d

07 Feb, 2024 1 commit

⚠

️ Raise `Exception` when trying to generate 0 tokens

⚠

️ (#28621) · abf8f54a

Daniel Korat authored Feb 07, 2024



* change warning to exception

* Update src/transformers/generation/utils.py
Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>

* validate `max_new_tokens` > 0 in `GenerationConfig`

* fix truncation test parameterization in `TextGenerationPipelineTests`

---------
Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>

abf8f54a