Commits · 5110e5748e7256bec1186c36d12a4ceaf018b879 · chenpangpang / transformers

16 Mar, 2023 3 commits

(#22204) · 5110e574

Yih-Dar authored Mar 16, 2023



* py38 + torch 2

* increment cache versions

---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

5110e574

LLaMA Implementation (#21955) · 0041be5b

Jason Phang authored Mar 16, 2023



* LLaMA

* sharding and docs

* tweak

* black

* inits

* ruff

* LLAMA_PRETRAINED_CONFIG_ARCHIVE_MAP

* init

* no checkpoint

* docs

* ruff

* type_vocab_size

* tokenizer fixes

* tokenizer fixes

* Update tokenization_llama.py

* Update tokenization_llama.py

* Update configuration_llama.py

* Update modeling_llama.py

* tokenizer add_bos by default

* licenses

* remove decoder

* norms and mlp

* rope overhaul

* tweaks

* black

* mention OPT implementation

* off-by-one naming

* typo

* fix

* tokenization fix and slicing bug

* padding config

* cleanup

* black

* update tests

* undo typo

* fix vocab caching logic

* ruff

* docbuilder

* attn fix from BlackSamorez

* initial feedback

* typo

* docs

* llama case

* llama case

* load checkpoint docs

* comment about tokenizer

* tokenizer defaults

* clear past_key_values if use_cache=False

* last tweaks

* last tweaks

* last tweaks

* last tweaks

---------
Co-authored-by: Stella Biderman <stellabiderman@gmail.com>

0041be5b

Update expected values in `MgpstrModelIntegrationTest` (#22195) · 52a57f7c
Yih-Dar authored Mar 16, 2023
```
Update values
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
52a57f7c

15 Mar, 2023 3 commits

Update BridgeTowerForContrastiveLearning (#22145) · 16121bae

Anahita Bhiwandiwalla authored Mar 15, 2023



* Use return_loss for BridgeTowerForContrastiveLearning, add example

* fix tests

* Update example in BridgeTowerForContrastiveLearning

* Update test_modeling_bridgetower.py

* update model output format

* minor update

* Update src/transformers/models/bridgetower/modeling_bridgetower.py

* make style

---------
Co-authored-by: Tiep Le <97980157+tileintel@users.noreply.github.com>
Co-authored-by: Tiep Le <tiep.le@intel.com>
Co-authored-by: Yih-Dar <2521628+ydshieh@users.noreply.github.com>
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

16121bae

Regression pipeline device (#22190) · 42ad693b
Sylvain Gugger authored Mar 15, 2023
```
* Fix regression in pipeline when device=-1 is passed

* Add regression test
```
42ad693b
Revert 22152 MaskedImageCompletionOutput changes (#22187) · 73768147
amyeroberts authored Mar 15, 2023
```
Revert changes
```
73768147

14 Mar, 2023 4 commits

to_pil - don't rescale if int and in range 0-255 (#22158) · c6318c37

amyeroberts authored Mar 14, 2023

* Don't rescale if in and in range 0-255

* Raise value error if int values too large

* Update tests/test_image_transforms.py

* Update tests/test_image_transforms.py

c6318c37

Create MaskedImageCompletionOutput and fix ViT docs (#22152) · 3b22bfbc
Alara Dirik authored Mar 14, 2023
```
* create MaskedImageCompletionOutput

* fix bugs

* fix bugs
```
3b22bfbc

Add ConvNeXT V2 (#21679) · cdddfbff

Alara Dirik authored Mar 14, 2023

* Add ConvNeXt V2 to transformers
* TF model is separated from the PR to fix issues

cdddfbff

Move `is_pipeline_test_to_skip` to specific model test classes (#21999) · 6c2ad00c

Yih-Dar authored Mar 14, 2023



* Move `is_pipeline_test_to_skip` to specific model test classes

---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

6c2ad00c

13 Mar, 2023 4 commits

[Safetensors] Add explicit flag to from pretrained (#22083) · f780557a

Patrick von Platen authored Mar 13, 2023



* [Safetensors] Add explicit  flag to from pretrained

* add test

* remove @

* Apply suggestions from code review
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

---------
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

f780557a

[`Whiper`] add `get_input_embeddings` to `WhisperForAudioClassification` (#22133) · d979cf6e

Younes Belkada authored Mar 13, 2023



* add `get_input_embeddings` to `WhisperForAudioClassification`

* add common tests

* fix another common test

* Update tests/models/whisper/test_modeling_whisper.py
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* fix style

---------
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

d979cf6e

[`Blip2`] skip accelerate test (#22124) · 6652e7da
Younes Belkada authored Mar 13, 2023
```
skip accelerate test
```
6652e7da

add new model of MGP-STR (#21418) · 102b5ff4

wangpeng authored Mar 13, 2023



* add new model of MGP-STR

* fix the check failings

* remove torch and numpy from mgp_tokenization

* remove unused import from modeling_mgp_str

* add test_processing_mgp_str

* rm test_processing_mgp_str.py

* add test_processing_mgp_str

* add test_processing_mgp_str

* add test_processing_mgp_str

* rm test_processing_mgp_str and add softmax outs to model

* rm test_processing_mgp_str and add softmax outs to model

* rewrite the code of mgp-str according to PR suggestions

* rewrite the code of mgp-str according to PR suggestions

* add new model of MGP-STR

* fix the check failings

* remove torch and numpy from mgp_tokenization

* remove unused import from modeling_mgp_str

* add test_processing_mgp_str

* rm test_processing_mgp_str.py

* add test_processing_mgp_str

* add test_processing_mgp_str

* add test_processing_mgp_str

* rm test_processing_mgp_str and add softmax outs to model

* rewrite the code of mgp-str according to PR suggestions

* rewrite the code of mgp-str according to PR suggestions

* remove representation_size from MGPSTRConfig

* reformat configuration_mgp_str.py

* format test_processor_mgp_str.py

* add test for tokenizer and complete model/processer test and model file

* rm Unnecessary tupple in modeling_mgp_str

* reduce hidden_size/layers/label_size in test_model

* add integration tests and change MGPSTR to Mgpstr

* add test for logit values

* reformat test model file

---------
Co-authored-by: yue kun <yuekun.wp@alibaba-inc.com>

102b5ff4

10 Mar, 2023 3 commits
- Revert "[GPT2] Propose fix for #21080" (#22093) · 2f320661
  Yih-Dar authored Mar 10, 2023
```
Revert "[GPT2] Propose fix for #21080 (#21853)" to avoid CI failure

This reverts commit a3fef89b.
```
  2f320661
- handle numpy inputs in whole word mask data collator (#22032) · 2f4cdd97
  Dean Wyatte authored Mar 10, 2023
  
  2f4cdd97
- [GPT2] Propose fix for #21080 (#21853) · a3fef89b
  Arthur authored Mar 10, 2023
```
* Make sure position ids are masked

* test that padded input produce the same results

* fix failing tests

* fixup

* fix batch test
```
  a3fef89b
09 Mar, 2023 3 commits

Skip 3 tests for `WhisperEncoderModelTest` (#22060) · ab81d31d
Yih-Dar authored Mar 09, 2023
```
* skip 3 tests

---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
ab81d31d
[deepspeed] offload + non-cpuadam optimizer exception (#22043) · ec24132b
Stas Bekman authored Mar 09, 2023
```
* [deepspeed] offload + non-cpuadam optimizer exception

* flip

* revert min version
```
ec24132b

Remove set_access_token usage + fail tests if FutureWarning (#22051) · 923110b7

Lucain authored Mar 09, 2023



* Remove set_access_token usage + fail tests if FutureWarning

* do not fail on FutureWarning in CI

---------
Co-authored-by: testbot <lucainp@hf.co>

923110b7

08 Mar, 2023 3 commits

Mark all `BridgeTower` tests slow for now (#22039) · 1cbac686
Yih-Dar authored Mar 08, 2023
```
* slow me

---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
1cbac686

[WIP] Add BridgeTowerForContrastiveLearning (#21964) · de81adf9

Anahita Bhiwandiwalla authored Mar 08, 2023



* Add BridgeTower for ITC

* Fix review feedback

* Rename BridgeTowerForITC, cleanup

* Fix style and quality

* implement tests

---------
Co-authored-by: Tiep Le <97980157+tileintel@users.noreply.github.com>
Co-authored-by: Tiep Le <tiep.le@intel.com>

de81adf9

Update `AudioClassificationPipelineTests::test_small_model_pt` for PT 2.0.0 (#22023) · dfe9a319
Yih-Dar authored Mar 08, 2023
```
fix
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
dfe9a319

07 Mar, 2023 10 commits

Update tiny model creation script and some others files (#22006) · b338414e

Yih-Dar authored Mar 07, 2023



* Update 1

* Update 2

* Update 3

* Update 4

* Update 5

* Update 6

* Update 7

* Update 8

* Update 9

* Update 10

---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

b338414e

[Time-Series] informer model (#21099) · 8abe4930

Eli Simhayev authored Mar 08, 2023

* added informer to gitignore

* added informer to gitignore

* WIP informer2020

* added checking that instantiate works

* added config using gluonTS by kashif

* WIP config

* adding informeConfig. need to remove FeatureEmbedder

* done InformerConfig, but need to change the names

* Done informer model init. working on enc-dec

* added things to address, after reading again enc-dec in the paper

* done modeling - checking initialization work

* added informer to gitignore

* WIP informer2020

* added checking that instantiate works

* added config using gluonTS by kashif

* WIP config

* adding informeConfig. need to remove FeatureEmbedder

* done InformerConfig, but need to change the names

* Done informer model init. working on enc-dec

* added things to address, after reading again enc-dec in the paper

* done modeling - checking initialization work

* moved enc-dec init to InformerEncoder/Decoder init

* added 'init_std' to config, now model init works!

* WIP conversion script, and added code sources

* WIP conversion script: loading original informer pth works

* WIP conversion script: change defaults in the config

* WIP conversion script: supporting Informer input embedding

* WIP conversion script: added parameters for the informer embed

* WIP conversion script: change dim_feedforward=2048

* WIP conversion script: remove unused args for loading checkpoint

* just cleaning up

* DataEmbedding removed, after thinking with Kashif

* working on forward pass

* WIP forward pass: trying to establish working batch for forward pass

* cleaning and finalizing

* adding HF names and docs

* init after cleaning works

* WIP in tests

* added docs for the informer specific args

* fix style

* undo change

* cleaning informer, now need to work only enc-dec

* initial enc-dec classes

* added encoder and decoder

* added todo

* add todos for conv_layers

* added decoder docs from vanilla

* added encoder docs from vanilla

* remove encoder decoder from the original informer

* removed AttentionLayer from the original paper

* removed TriangularCausalMask, same as decoder_attention_mask

* initial sparse attention

* use conv_layers

* fixed test_config test

* fix parenthesis when itearting zip(layers, conv_layers)

* error found in prob attention, added sizes as comments

* fix sizes

* added proposal for q_reduce indexing, and remove unused

* WIP ProbMask, and changed factor=2 for testing

* remove unused libs for this PR for creating the env

* fix checking the attn_weights.size() after bmm

* Q_reduce: changed from torch.gather to simple slicing

* WIP calculate final attn_output

* finish adding v_aggregated, attn_output ready

* changed tgt_len to u in attention_mask, need to fix the size error

* comment attention_mask for encoder, and fix if cond for v_agg

* added ProbMask support (wip), removed old original code

* finished ProbMask 😃



* Revert "remove unused libs for this PR for creating the env"

This reverts commit 11a081e09e92771e51a5d2758d53a9afb59547f0.

* fixes

* make style

* fix initial tests

* fix more tests

* dry

* make style

* remove unused files

* style

* added integration tests

* fix num_static_real_features

* fix header

* remove unused function

* fix example

* fix docs

* Update src/transformers/models/informer/configuration_informer.py
Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>

* Update src/transformers/models/informer/modeling_informer.py
Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>

* Update src/transformers/models/informer/configuration_informer.py
Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>

* Update src/transformers/models/informer/configuration_informer.py
Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>

* Update src/transformers/models/informer/configuration_informer.py
Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>

* Update src/transformers/models/informer/configuration_informer.py
Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>

* fixes for reviewer

* use prediction_length from model

* fix style

* fixed informer.mdx

* added to index

* updated readme

* undo

* make fix-copies

* typo

* fix copy

* added Informer to toctree

* in order

* fixed comments

* remove unneeded new lines in docs

* make static real and cat optional

* fix use of distil conv layers

* fixed integration test

* added checkpoint for convlayer

* make fix-copies

* updated from time series model

* make fix-copies

* copy decoder

* fix unit tests

* updated scaling config

* fix integration tests

* IGNORE_NON_TESTED

* IGNORE_NON_AUTO_CONFIGURED

* IGNORE_NON_AUTO_CONFIGURED

* updated check configs

* fix formatting

* undo change from time series

* prediction_length should not be None

* aliign with the blog: prettify ProbSparse and change attention_factor  to sampling_factor

* make style

* make fix-copies

* niels CR: update contributed by

* niels CR: update configuration_informer.py
Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>

* niels CR: update kashif -> huggingface
Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>

* niels CR: `sampling_factor` only relevant when `attention_type`=prob

* make style

* fixed U_part: added multiplication by `L_Q`

* fixed bug: remove `is not None` from `if config.distil`

* fixed test: `decoder_seq_length` to `encoder_seq_length` in cross_attentions check

* fix integration tests

* updated model hub

* do not shift as in training

* undo

* fix make-copies

* make fix-copies

* added `if prediction_length is None`

* changed `ProbSparseAttention` to `InformerProbSparseAttention`

* changed `V_sum` -> `v_mean_dim_time`

* changed `ConvLayer` to `InformerConvLayer` and fixed `super()`

* TimeSeriesTansformer->Informer in decoder's Copied from

* more descriptive in ProbSparse

* make style

* fix coped from

* Revert "added `if prediction_length is None`"

This reverts commit b4cbddfa05e3bd739b79569cd3c3b89e316f2451.

* fixed indent

* use InformerSinusoidalPositionalEmbedding

* make fix-style

* fix from #21860

* fix name

* make fix-copies

* use time series utils

* fix dec num_heads

* docstring

* added time series util doc

* _import_structure

* formatting

* changes from review

* make style

* fix docs

* fix doc

* removed NegativeLogLikelihood

---------
Co-authored-by: Kashif Rasul <kashif.rasul@gmail.com>
Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>

8abe4930

[DETR and friends] Remove is_timm_available (#21814) · dde718e7

NielsRogge authored Mar 07, 2023



* First draft

* Fix to_dict

* Improve conversion script

* Update config

* Remove timm dependency

* Fix dummies

* Fix typo, add integration test

* Upload 101 model as well

* Remove timm dummies

* Fix style

---------
Co-authored-by: Niels Rogge <nielsrogge@Nielss-MacBook-Pro.local>

dde718e7

[TF] Fix creating a PR while pushing in TF framework (#21968) · 2156662d

Arthur authored Mar 07, 2023

* add create pr arg

* style

* add test

* ficup

* update test

* last nit fix typo

* add `is_pt_tf_cross_test` marker for the tsts

2156662d

[Whisper] Add model for audio classification (#21754) · 7c393181

Sanchit Gandhi authored Mar 07, 2023

* [Whisper] Add model for audio classification

* make fix-copies

* add to docs

* add docstring

* empty returns

* add code example

* switch to fleurs

* stick everything on one line

7c393181

Skip `test_multi_gpu_data_parallel_forward` for some model tests (#21991) · 9402788b

Yih-Dar authored Mar 07, 2023



skip test_multi_gpu_data_parallel_forward for some model tests
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

9402788b

[DETR, YOLOS] Fix device bug (#21974) · 95408e99
NielsRogge authored Mar 07, 2023
```
* Fix integration test

* Add test

* Add test
```
95408e99
Fix MinNewTokensLengthLogitsProcessor when used with a list of eos tokens (#21959) · eec46b4f
Elad Segal authored Mar 07, 2023
```
* Fix MinNewTokensLengthLogitsProcessor when used with a list of eos tokens

* fix docs

* Empty commit

* formatting
```
eec46b4f
Add check before int casting for PIL conversion (#21969) · 4063fd9c
amyeroberts authored Mar 07, 2023
```
* Add check before int casting for PIL conversion

* Line length

* Tidier logic
```
4063fd9c

Update `Jukebox` tests (#21984) · 5b28b783

Yih-Dar authored Mar 07, 2023



* update expected values for jukebox

* update expected values for jukebox

* update expected values for jukebox

* update expected values for jukebox

* update expected values for jukebox

---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

5b28b783

06 Mar, 2023 3 commits
- Update expected values for `test_xglm_sample` (#21975) · f2a2616b
  Yih-Dar authored Mar 06, 2023
```
update expected values for xglm
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
  f2a2616b
- Use larger atol in `torch.allclose` for some tests (#21966) · 9474abdf
  Yih-Dar authored Mar 06, 2023
```
Use larger atol
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
  9474abdf
- Update expected values in `XLMProphetNetModelIntegrationTest` (#21957) · fcf81341
  Yih-Dar authored Mar 06, 2023
```
update values
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
  fcf81341
03 Mar, 2023 3 commits

[CLAP] Support batched inputs for CLAP. Fixes pipeline issues (#21931) · 718e9d77

Arthur authored Mar 03, 2023

* fix pipeline

* fix feature_extraction clap

* you can now batch the `is_longer` attribute

* add tests

* fixup

* add expected scores

* comment on is_longert

718e9d77

Fix `AlignModelTest` tests (#21923) · d4306dae

Yih-Dar authored Mar 03, 2023



* fix

* fix

---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

d4306dae

Update `model_split_percents` for `WhisperModelTest` (#21922) · fa9d2ad7
Yih-Dar authored Mar 03, 2023
```
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
fa9d2ad7

02 Mar, 2023 1 commit

Avoid modeling tests run in pipeline CI jobs (#21911) · 9f5bfe1b

Yih-Dar authored Mar 02, 2023



* rework is_pipeline_test

* bring back 3 tests

---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

9f5bfe1b