Commits · f243a5ec0d7a5a30e8bf76f1750ed0a959d47fa9 · chenpangpang / transformers

12 Apr, 2021 9 commits

Sagemaker test docs update for framework upgrade (#11206) · f243a5ec
Philipp Schmid authored Apr 13, 2021
```
* increased train_runtime for model parallelism

* added documentation for framework upgrade
```
f243a5ec
Import torch.utils.checkpoint in ProphetNet (#11214) · 74d7c24d
Lysandre Debut authored Apr 12, 2021

74d7c24d
Replaced `which` with `who` (#11183) · 38a10c6b
cronoik authored Apr 13, 2021

38a10c6b

NielsRogge authored Apr 13, 2021

* First draft of deit

* More improvements

* Remove DeiTTokenizerFast from init

* Conversion script works

* Add DeiT to ViT conversion script

* Add tests, add head model, add support for deit in vit conversion script

* Update model checkpoint names

* Update image_mean and image_std, set resample to bicubic

* Improve docs

* Docs improvements

* Add DeiTForImageClassificationWithTeacher to init

* Address comments by @sgugger

* Improve feature extractors

* Make fix-copies

* Minor fixes

* Address comments by @patil-suraj

* All models uploaded

* Fix tests

* Remove labels argument from DeiTForImageClassificationWithTeacher

* Fix-copies, style and quality

* Fix tests

* Fix typo

* Multiple docs improvements

* More docs fixes

9f126097

Fix typo (#11188) · cb251ba6
Takuya Makino authored Apr 13, 2021

cb251ba6

Added documentation for data collator. (#10941) · 0c6fcd30

fghuman authored Apr 12, 2021



* Added documentation for data collator.

* Update docs/source/data_collator.rst
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Added documentation for data collator.

* Added documentation for the data collator.

* Merge branch 'doc_DataCollator' of C:\Users\mahii\PycharmProjects\transformers with conflicts.

* Update documentation for the data collator.

* Update documentation for the data collator.
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>
Co-authored-by: Amna <A.A.Ahmad@student.tudelft.nl>

0c6fcd30

model_path should be ignored as the checkpoint path (#11157) · ef102c48

Masatoshi TSUCHIYA authored Apr 12, 2021

* model_path is refered as the path of the trainer, and should be ignored as the checkpoint path.

* Improved according to Sgugger's comment.

ef102c48

Fix style · 623cd6ae
Sylvain Gugger authored Apr 12, 2021

623cd6ae
Minor typos fixed (#11182) · a99f7f5c
cronoik authored Apr 12, 2021

a99f7f5c

09 Apr, 2021 13 commits
- Reactivate Megatron tests an use less workers · 26212c14
  Sylvain Gugger authored Apr 09, 2021
  
  26212c14
- Fix Typo · 716120cb
  Lysandre authored Apr 09, 2021
  
  716120cb
- added json dump and extraction of train run time (#11167) · 6f90c29e
  Philipp Schmid authored Apr 09, 2021
```
* added json dump and extraction of train run time

* make style happy
```
  6f90c29e
- [examples run_clm] fix _LazyModule hasher error (#11168) · 07f0bb69
  Stas Bekman authored Apr 09, 2021
```
* fix _LazyModule hasher error

* reword
```
  07f0bb69
- [examples/translation] support mBART-50 and M2M100 fine-tuning (#11170) · c161dd56
  Suraj Patil authored Apr 09, 2021
```
* keep a list of multilingual tokenizers

* add forced_bos_token argument
```
  c161dd56
- Add a special tokenizer for CPM model (#11068) · fb41f9f5
  Kevin Canwen Xu authored Apr 10, 2021
```
* Add a special tokenizer for CPM model

* make style

* fix

* Add docs

* styles

* cpm doc

* fix ci

* fix the overview

* add test

* make style

* typo

* Custom tokenizer flag

* Add REAMDE.md
Co-authored-by: Lysandre <lysandre.debut@reseau.eseo.fr>
```
  fb41f9f5
- Make `get_special_tokens_mask` consider all tokens (#11163) · 45fc8c79
  Sylvain Gugger authored Apr 09, 2021
  
  45fc8c79
- Update README.md (#11161) · 60607465
  Saviour Owolabi authored Apr 09, 2021
```
Corrected a typo ('Downlowd' to 'Download')
```
  60607465
- Fix LogitsProcessor documentation (#11130) · b9b60c16
  Keisuke Hirota authored Apr 09, 2021
```
* Change duplicated LogitsProcessor to LogitsWarper in LogitsProcessorList document

* Write more detailed information about LogitsProcessor's scores argument

* apply suggestion from review

* style
Co-authored-by: Suraj Patil <surajp815@gmail.com>
```
  b9b60c16
- [Community notebooks] Add Wav2Vec notebook for creating captions for YT Clips (#11142) · 8b78a32b
  Niklas Muennighoff authored Apr 09, 2021
```
* Add Wav2Vec Inference notebook

* Update docs/source/community.md
Co-authored-by: Suraj Patil <surajp815@gmail.com>
```
  8b78a32b
- typo (#11152) · 0311ba21
  Stas Bekman authored Apr 08, 2021
```
* typo

* style
```
  0311ba21
- Merge branch 'master' of github.com:huggingface/transformers · 269c9638
  Sylvain Gugger authored Apr 08, 2021
  
  269c9638
- Skip Megatron tests for now · d31c7b10
  Sylvain Gugger authored Apr 08, 2021
  
  d31c7b10
08 Apr, 2021 14 commits

[setup] make fairscale and deepspeed setup extras (#11151) · c2e0fd52

Stas Bekman authored Apr 08, 2021



* make fairscale and deepspeed setup extras

* fix default

* Apply suggestions from code review
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* no reason not to ask for the good version

* update the CIs
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

c2e0fd52

Add support for multiple models for one config in auto classes (#11150) · ba8b1f47
Sylvain Gugger authored Apr 08, 2021
```
* Add support for multiple models for one config in auto classes

* Use get_values everywhere

* Prettier doc
```
ba8b1f47
[setup] extras[docs] must include 'all' (#11148) · 97ccf67b
Stas Bekman authored Apr 08, 2021
```
* extras[doc] must include 'all'

* fix

* better

* regroup
```
97ccf67b

[tests] relocate core integration tests (#11146) · 66446909

Stas Bekman authored Apr 08, 2021

* relocate core integration tests

* add sys.path context manager

* cleanup

* try

* try2

* fix path

* doc

* style

* add dep

* add 2 more deps

66446909

Run mlm pad to multiple for fp16 (#11128) · 6c40e497
Andrea Cappelli authored Apr 08, 2021
```
* Add mlm collator pad to multiple option (#10627)

* Use padding to 8x in run mlm (#10627)
```
6c40e497
Don't duplicate logs in TensorBoard and handle --use_env (#11141) · dfed4ec2
Sylvain Gugger authored Apr 08, 2021

dfed4ec2
Updates SageMaker docs for updating DLCs (#11140) · 9c9b8e70
Philipp Schmid authored Apr 08, 2021

9c9b8e70
Add fairscale and deepspeed back to the CI (#11147) · ba2cf5f9
Lysandre Debut authored Apr 08, 2021
```
* Add fairscale and deepspeed back to the CI

* Add deepspeed to single GPU tests
```
ba2cf5f9
[trainer] solve "scheduler before optimizer step" warning (#11144) · 1ed24afe
Stas Bekman authored Apr 08, 2021
```
* solve "scheduler before optimizer step" warning

* style

* correct the state evaluation test
```
1ed24afe

Add nvidia megatron models (#10911) · 02ec02d6

Julien Demouth authored Apr 08, 2021



* Add support for NVIDIA Megatron models

* Add support for NVIDIA Megatron GPT2 and BERT

Add the megatron_gpt2 model. That model reuses the existing GPT2 model. This
commit includes a script to convert a Megatron-GPT2 checkpoint downloaded
from NVIDIA GPU Cloud. See examples/megatron-models/README.md for details.

Add the megatron_bert model. That model is implemented as a modification of
the existing BERT model in Transformers. This commit includes a script to
convert a Megatron-BERT checkpoint downloaded from NVIDIA GPU Cloud. See
examples/megatron-models/README.md for details.

* Update src/transformers/models/megatron_bert/configuration_megatron_bert.py
Co-authored-by: Lysandre Debut <lysandre@huggingface.co>

* Update src/transformers/models/megatron_bert/configuration_megatron_bert.py
Co-authored-by: Lysandre Debut <lysandre@huggingface.co>

* Update src/transformers/models/megatron_bert/configuration_megatron_bert.py
Co-authored-by: Lysandre Debut <lysandre@huggingface.co>

* Remove model.half in tests + add "# Copied ..."

Remove the model.half() instruction which makes tests fail on the CPU.

Add a comment "# Copied ..." before many classes in the model to enable automatic
tracking in CI between the new Megatron classes and the original Bert ones.

* Fix issues

* Fix Flax/TF tests

* Fix copyright

* Update src/transformers/models/megatron_bert/configuration_megatron_bert.py
Co-authored-by: Lysandre Debut <lysandre@huggingface.co>

* Update src/transformers/models/megatron_bert/configuration_megatron_bert.py
Co-authored-by: Lysandre Debut <lysandre@huggingface.co>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Lysandre Debut <lysandre@huggingface.co>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Lysandre Debut <lysandre@huggingface.co>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Lysandre Debut <lysandre@huggingface.co>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Lysandre Debut <lysandre@huggingface.co>

* Update docs/source/model_doc/megatron_bert.rst
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update docs/source/model_doc/megatron_gpt2.rst
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/__init__.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_gpt2/convert_megatron_gpt2_checkpoint.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_gpt2/convert_megatron_gpt2_checkpoint.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_gpt2/convert_megatron_gpt2_checkpoint.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/convert_megatron_bert_checkpoint.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/convert_megatron_bert_checkpoint.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/convert_megatron_bert_checkpoint.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/models/megatron_bert/modeling_megatron_bert.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Resolve most of 'sgugger' comments

* Fix conversion issue + Run make fix-copies/quality/docs

* Apply suggestions from code review

* Causal LM & merge

* Fix init

* Add CausalLM to last auto class
Co-authored-by: Julien Demouth <jdemouth@nvidia.com>
Co-authored-by: Lysandre Debut <lysandre@huggingface.co>
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>
Co-authored-by: Lysandre <lysandre.debut@reseau.eseo.fr>

02ec02d6

[DeepSpeed] ZeRO Stage 3 (#10753) · c6d66484

Stas Bekman authored Apr 08, 2021



* synced gpus

* fix

* fix

* need to use t5-small for quality tests

* notes

* complete merge

* fix a disappearing std stream problem

* start zero3 tests

* wip

* tune params

* sorting out the pre-trained model loading

* reworking generate loop wip

* wip

* style

* fix tests

* split the tests

* refactor tests

* wip

* parameterized

* fix

* workout the resume from non-ds checkpoint pass + test

* cleanup

* remove no longer needed code

* split getter/setter functions

* complete the docs

* suggestions

* gpus and their compute capabilities link

* Apply suggestions from code review
Co-authored-by: Lysandre Debut <lysandre@huggingface.co>

* style

* remove invalid paramgd

* automatically configure zero3 params that rely on hidden size

* make _get_resized_embeddings zero3-aware

* add test exercising resize_token_embeddings()

* add docstring
Co-authored-by: Lysandre Debut <lysandre@huggingface.co>

c6d66484

[run_clm] clarify why we get the tokenizer warning on long input (#11145) · acc851e1

Stas Bekman authored Apr 08, 2021



* clarify why we get the warning here

* Update examples/language-modeling/run_clm.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* wording

* style
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

acc851e1

Typo fix of the name of BertLMHeadModel in BERT doc (#11133) · 5bf5d50c
Yusuke Mori authored Apr 08, 2021

5bf5d50c

Fix typing error in Trainer class (prediction_step) (#11138) · f8e90d6f

Jannis Born authored Apr 08, 2021

* fix: docstrings in prediction_step

* ci: Satisfy line length requirements

* ci: character length requirements

f8e90d6f

07 Apr, 2021 4 commits
- Fix and refactor check_repo (#11127) · ffe07617
  Sylvain Gugger authored Apr 07, 2021
  
  ffe07617
- Adds use_auth_token with pipelines (#11123) · 3fd7eee1
  Philipp Schmid authored Apr 07, 2021
```
* added model_kwargs to infer_framework_from_model

* added model_kwargs to tokenizer

* added use_auth_token as named parameter

* added dynamic get for use_auth_token
```
  3fd7eee1
- [versions] handle version requirement ranges (#11110) · 1c151283
  Stas Bekman authored Apr 07, 2021
```
* handle version requirement ranges

* add mixed requirement test

* cleanup
```
  1c151283
- fix tests (#11109) · 7442801d
  Vasudev Gupta authored Apr 07, 2021
  
  7442801d