Commits · 3857f2b4e34912c942694489c2b667d9476e55f5 · chenpangpang / transformers

07 Jun, 2021 2 commits
- fix deberta 2 tokenizer integration test (#12017) · 3857f2b4
  Philip May authored Jun 07, 2021
  
  3857f2b4
- Fixed Typo in modeling_bart.py (#12035) · 20b6f3b8
  Shiva Pundir authored Jun 07, 2021
```
* Fixed Typo in modeling_bart.py - Issue #11895

* Fixed Typo in modeling_bart.py
```
  20b6f3b8
04 Jun, 2021 2 commits

[TrainerArguments] format and sort __repr__, add __str__ (#12018) · 1f335aef
Stas Bekman authored Jun 04, 2021
```
* format and sort __repr__, add __str__

* typo

* use __str__ directly

* alias __repr__ = __str__
```
1f335aef

[Deepspeed] Assert on mismatches between ds and hf args (#12021) · 2c73b930

Stas Bekman authored Jun 04, 2021



* wip

* add mismatch validation + test

* renames

* Update docs/source/main_classes/deepspeed.rst
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* renames
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

2c73b930

03 Jun, 2021 2 commits

[Flax] Refactor MLM (#12013) · 242ec31a

Patrick von Platen authored Jun 03, 2021



* fix_torch_device_generate_test

* remove @

* finish refactor
Co-authored-by: Patrick von Platen <patrick@huggingface.co>

242ec31a

Fix weight decay masking in `run_flax_glue.py` (#11964) · 4674061b

Nicholas Vadivelu authored Jun 03, 2021



* Fix weight decay masking in `run_flax_glue.py`

Issues with the previous implementation:
- The `dict` from `traverse_util.flatten_dict` has keys which are tuples of strings, not one long string with the path separated by periods.
- `optax.masked` applies the transformation wherever the mask is True, so the masks are flipped.
- Flax's LayerNorm calls the scale parameter `scale` not `weight`

* Fix formatting with black

* adapt results
Co-authored-by: Patrick von Platen <patrick@huggingface.co>

4674061b

02 Jun, 2021 8 commits

[deepspeed] add nvme test skip rule (#11997) · 61c50634
Stas Bekman authored Jun 02, 2021
```
* add nvme skip rule

* fix
```
61c50634
[deepspeed] Move code and doc into standalone files (#11984) · 640318be
Stas Bekman authored Jun 02, 2021
```
* move code and docs

* style

* moved

* restore
```
640318be

Update return introduction (#11976) · d6d747cb

Kou Yong Kang authored Jun 03, 2021

Make it clear that the `forward` method now returns a dict instead of tuple.

Fix style

d6d747cb

[docs] fix xref to `PreTrainedModel.generate` (#11049) · d406a272
Stas Bekman authored Jun 02, 2021
```
* fix xref to generate

* do the same for search methods

* style

* style
```
d406a272
Fix examples (#11990) · 123b597f
Gunjan Chhablani authored Jun 02, 2021

123b597f

VisualBERT (#10534) · 88ca6a23

Gunjan Chhablani authored Jun 02, 2021



* Init VisualBERT

* Add cookie-cutter, Config, and Embeddings

* Add preliminary Model

* Add Bert analogous classes

* Add basic code for NLVR, VQA, Flickr

* Update Init

* Fix VisualBert Downstream Models

* Rename classifier to cls

* Comment position_ids buffer

* Remove sentence image predictor output

* Update output dicts

* Remove unnecessary files

* Fix Auto Modeling

* Fix transformers init

* Add conversion script

* Add conversion script

* Fix docs

* Update visualbert modelling

* Update configuration

* Style fixes

* Add model and integration tests

* Add all tests

* Update model mapping

* Add simple detector from original repository

* Update docs and configs

* Fix style

* Fix style

* Update docs

* Fix style

* Fix import issues in style

* Fix style

* Add changes from review

* Fix style

* Fix style

* Update docs

* Fix style

* Fix style

* Update docs/source/model_doc/visual_bert.rst
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/visual_bert/modeling_visual_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update tests/test_modeling_visual_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/visual_bert/modeling_visual_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/visual_bert/modeling_visual_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/visual_bert/modeling_visual_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Add changes from review

* Remove convert run script

* Add changes from review

* Update src/transformers/models/visual_bert/modeling_visual_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/visual_bert/modeling_visual_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/visual_bert/modeling_visual_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/visual_bert/modeling_visual_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/visual_bert/modeling_visual_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Add changes from review

* Add changes from review

* Add visual embedding example in docs

* Fix "copied from" comments

* Add changes from review

* Fix error, style, checkpoints

* Update docs

* Fix integration tests

* Fix style
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

88ca6a23

[RAG] Fix rag from pretrained question encoder generator behavior (#11962) · 43f46aa7
Patrick von Platen authored Jun 02, 2021
```
* fix_torch_device_generate_test

* remove @

* fix rag from pretrained loading

* add test

* uplaod

* finish
```
43f46aa7

Bump urllib3 from 1.25.8 to 1.26.5 in /examples/research_projects/lxmert (#11983) · 6db3a87d

dependabot[bot] authored Jun 02, 2021

Bumps [urllib3](https://github.com/urllib3/urllib3) from 1.25.8 to 1.26.5.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/1.25.8...1.26.5

)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-type: direct:production
...
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

6db3a87d

01 Jun, 2021 15 commits

[Trainer] add train loss and flops metrics reports (#11980) · 4ba203d9

Stas Bekman authored Jun 01, 2021

* add train loss and flops metrics reports

* consistency

* add train_loss to skip keys

* restore on_train_end call timing

4ba203d9

[DeepSpeed] decouple `DeepSpeedConfigHF` from `Trainer` (#11966) · 7ec596ec

Stas Bekman authored Jun 01, 2021



* decouple DeepSpeedConfigHF from Trainer

* add LoggingLevel ctx manager; add new test

* cleanup

* add docs

* Apply suggestions from code review
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* implemented suggested renames

* formatter workaround
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

7ec596ec

Typo in usage example, changed to device instead of torch_device (#11979) · 1c3ab3e5
Alberto Villa authored Jun 01, 2021

1c3ab3e5

ByT5 model (#11971) · 47a98fc4

Patrick von Platen authored Jun 01, 2021



* allow tf to use uneven num of layers

* add tokenizer

* finish docs

* finish docs

* Apply suggestions from code review

* include in index

* finish

* Update docs/source/model_doc/byt5.rst
Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>

* apply sylvais suggestions

* make style
Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>

47a98fc4

typo correction (#11973) · 1eb58b45
Jeoung-Minju authored Jun 02, 2021
```
* typo correction

* type corrections
```
1eb58b45
[deepspeed] docs (#11940) · 79712e7e
Stas Bekman authored Jun 01, 2021
```
* deepspeed docs

* cleanup

* cleanup
```
79712e7e
Run the integration tests on schedule tests instead of master tests · 985d7088
Lysandre authored Jun 01, 2021

985d7088

Neptune.ai integration (#11937) · 9996558b

Volodymyr Byno authored Jun 01, 2021

An option that turns on neptune.ai logging
--report_to 'neptune'

Additional ENV variables:
	NEPTUNE_PROJECT
	NEPTUNE_API_TOKEN
	NEPTUNE_RUN_NAME (optional)
	NEPTUNE_STOP_TIMEOUT (optional)

9996558b

Authorize args when instantiating an AutoModel (#11956) · ae6ce28f
Lysandre Debut authored Jun 01, 2021

ae6ce28f

Add regression tests for slow sentencepiece tokenizers. (#11737) · fcad8018

Philip May authored Jun 01, 2021

* add test_vocab_size for sentencepiece tok.

* add test_get_vocab for sentencepiece tok.

* add test_convert_token_and_id for sentencepiece tok.

* add test_tokenize_and_convert_tokens_to_string for all tok.

* improve test_tokenize_and_convert_tokens_to_string for sp. tok.

* add common tokenizer integration tests
- for albert
- for barthez

* add tokenizer integration tests to bert gen.

* add most tokenizer integration tests

* fix camembert tokenizer integration test

* add tokenizer integration test to marian

* add tokenizer integration test to reformer

* add typing and doc to tokenizer_integration_test_util

* fix tokenizer integration test of reformer

* improve test_sentencepiece_tokenize_and_convert_tokens_to_string

* empty commit to trigger CI

* fix tokenizer integration test of reformer

* remove code not needed anymore

* empty commit to trigger CI

* empty commit to trigger CI

fcad8018

reinitialize wandb config for each hyperparameter search run (#11945) · c3d958b2
Josh Tanner authored Jun 01, 2021

c3d958b2

bugfixes training_args.py (#11922) · 99dbbdb9

Riccardo Bassani authored Jun 01, 2021

modified according to:
https://pytorch.org/xla/release/1.8.1/_modules/torch_xla/core/xla_model.html

99dbbdb9

modify qa-trainer (#11872) · 7e73601f
Fan Zhang authored Jun 01, 2021
```
* modify qa-trainer

* fix flax model
```
7e73601f

RAG-2nd2end-revamp (#11893) · 9ec0f01b

Shamane Siri authored Jun 01, 2021



* initial

* code quality test

* code quality

* added test functions in test_modeling_rag.py and test_retrieval_rag.py to test end2end retreiver

* minor change in test_modeling_rag

* fixed tests

* Update examples/research_projects/rag-end2end-retriever/README.md

typo corrected as suggested by lhoestq
Co-authored-by: Quentin Lhoest <42851186+lhoestq@users.noreply.github.com>

* Update examples/research_projects/rag-end2end-retriever/finetune_rag.py

type change suggested by lhoestq
Co-authored-by: Quentin Lhoest <42851186+lhoestq@users.noreply.github.com>

* Update src/transformers/models/rag/retrieval_rag.py

Adding this change as mentioned by lhoestq.
Co-authored-by: Quentin Lhoest <42851186+lhoestq@users.noreply.github.com>

* completed the minor changes suggested by the reviewers
Co-authored-by: Quentin Lhoest <42851186+lhoestq@users.noreply.github.com>

9ec0f01b

Add FlaxCLIP (#11883) · ad25fd62

Suraj Patil authored Jun 01, 2021

* add flax CLIP

* default input_shape

* add tests

* fix test

* fix name

* fix docs

* fix shapes

* attend at least 1 token

* flax conv to torch conv

* return floats

* fix equivalence tests

* fix import

* return attention_weights and update tests

* fix dosctrings

* address patricks comments

* input_shape arg

* add tests for get_image_features and get_text_features methods

* fix tests

ad25fd62

31 May, 2021 4 commits
- Add MT5ForConditionalGeneration as supported arch. to summarization README (#11961) · cfca638a
  Philip May authored May 31, 2021
```
* Add MT5ForConditionalGeneration as supported arch.

* Update README.md
```
  cfca638a
- Remove redundant `nn.log_softmax` in `run_flax_glue.py` (#11920) · 1ab147d6
  Nicholas Vadivelu authored May 31, 2021
```
* Remove redundant `nn.log_softmax` in `run_flax_glue.py`

`optax.softmax_cross_entropy` expects unnormalized logits, and so it already calls `nn.log_softmax`, so I believe it is not needed here. `nn.log_softmax` is idempotent so mathematically it shouldn't have made a difference.

* Remove unused 'flax.linen' import
```
  1ab147d6
- fix assert (#11935) · fb60c309
  Philip May authored May 31, 2021
  
  fb60c309
- Remove `datasets` submodule · 04a9709c
  Lysandre authored May 31, 2021
  
  04a9709c
28 May, 2021 3 commits

Test optuna and ray (#11924) · 8d171628
Lysandre Debut authored May 28, 2021

8d171628

[Flax] Return Attention from BERT, ELECTRA, RoBERTa and GPT2 (#11918) · af1a10bf

Jayendra authored May 28, 2021



* Added logic to return attention from flax-bert model and added test cases to check that

* Added new line at the end of file to test_modeling_flax_common.py

* fixing code style

* Fixing Roberta and Elextra models too from cpoying bert

* Added temporary hack to not run test_attention_outputs for FlaxGPT2

* Returning attention weights from GPT2 and changed the tests accordingly.

* last fixes

* bump flax dependency
Co-authored-by: jayendra <jayendra@infocusp.in>
Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>

af1a10bf

Added Sequence Classification class in GPTNeo (#11906) · e1205e47
Bhadresh Savani authored May 28, 2021
```
* seq classification changes

* fix tests
```
e1205e47

27 May, 2021 3 commits

Adding new argument `max_new_tokens` for generate. (#11476) · 80d712fa

Nicolas Patry authored May 27, 2021

* Adding new argument `max_new_tokens` for generate.

This is a proposal to add a new argument `max_new_tokens` to `generate`.
This include a `MaxNewTokensCriteria` that enables callers that don't
know about the token length ahead (like pipelines callers) to manage
more easily the length of their generated output.

* Adding a test for the user warning when both`max_length` and
`max_new_tokens` are used together.

* Removed redundant `no_grad`.

80d712fa

Update deepspeed config to reflect hyperparameter search parameters (#11896) · 2dd6fb25
Josh Tanner authored May 27, 2021
```
* rebuild deepspeed config for hyperparameter search

* reformat code to fix style issues
```
2dd6fb25
Add Emotion Speech Noteboook (#11900) · 42fe0dc2
Patrick von Platen authored May 27, 2021

42fe0dc2

26 May, 2021 1 commit

Flax Generate (#11777) · 996a315e

Patrick von Platen authored May 27, 2021



* fix_torch_device_generate_test

* remove @

* add

* indexing

* correct a couple of tests

* fix tests

* add logits processor

* finish top_k, top_p, temp

* add docs

* correct flax prng key default

* improve generate

* add generation docs

* add docs

* make style

* revert model outputs change

* make style

* correct typo

* fix tests

* fix slow test

* add raise

* finish generation
Co-authored-by: Patrick von Platen <patrick@huggingface.co>

996a315e