Commits · 8bf7312654d40cea1a399e86f9fe8e39e1ea3a1e · chenpangpang / transformers

07 May, 2020 17 commits

Add AlbertForPreTraining and TFAlbertForPreTraining models. (#4057) · 8bf73126

Jared T Nielsen authored May 07, 2020



* Add AlbertForPreTraining and TFAlbertForPreTraining models.

* PyTorch conversion

* TensorFlow conversion

* style
Co-authored-by: Lysandre <lysandre.debut@reseau.eseo.fr>

8bf73126

[doc] Fix broken links + remove crazy big notebook · c99fe038
Julien Chaumond authored May 07, 2020

c99fe038
Create README.md (#4202) · 66113bd6
Savaş Yıldırım authored May 08, 2020

66113bd6
[examples] Add column for pytorch-lightning support · 6669915b
Julien Chaumond authored May 07, 2020

6669915b
Examples readme.md (#4215) · 612fa1b1
Julien Chaumond authored May 07, 2020
```
* README

* Update README.md
```
612fa1b1
Pin isort and tf <= 2.1.0 · 2e578243
Lysandre authored May 07, 2020

2e578243
Release: v2.9.0 · e7cfc1a3
Lysandre authored May 07, 2020

e7cfc1a3

BIG Reorganize examples (#4213) · 0ae96ff8

Julien Chaumond authored May 07, 2020

* Created using Colaboratory

* [examples] reorganize files

* remove run_tpu_glue.py as superseded by TPU support in Trainer

* Bugfix: int, not tuple

* move files around

0ae96ff8

[Trainer] Ability to specify optimizer/scheduler at init · cafa6a9e
Julien Chaumond authored May 07, 2020
```
cc @patrickvonplaten @thomwolf
```
cafa6a9e
Use with_extension to change the extension (#4203) · e4fd5e39
Bram Vanroy authored May 07, 2020
```
As per https://github.com/huggingface/transformers/pull/3934#discussion_r421307659
```
e4fd5e39

Tpu trainer (#4146) · ebf80e2e

Lysandre Debut authored May 07, 2020



* wip

* wip

* a last wip

* Better logging when using TPUs

* Correct argument name

* Tests

* fix

* Metrics in evaluation

* Update src/transformers/training_args.py

* [tpu] Use launcher script instead

* [tpu] lots of tweaks

* Fix formatting
Co-authored-by: Julien Chaumond <chaumond@gmail.com>

ebf80e2e

Ensure fast tokenizer can construct tensor without pad token if only one... · 026097b9
Funtowicz Morgan authored May 07, 2020
```
Ensure fast tokenizer can construct tensor without pad token if only one sample is provided. (#4201)
```
026097b9

Rewritten batch support in pipelines. (#4154) · 0a6cbea0

Funtowicz Morgan authored May 07, 2020



* Rewritten batch support in pipelines.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Fix imports sorting 🔧

Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Set pad_to_max_length=True by default on Pipeline.

* Set pad_to_max_length=False for generation pipelines.

Most of generation models doesn't have padding token.

* Address @joeddav review comment: Uniformized *args.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Address @joeddav review comment: Uniformized *args (second).
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

0a6cbea0

fix examples (#4192) · 99d1a694
Patrick von Platen authored May 07, 2020

99d1a694
[Reformer] Fix example and error message (#4191) · 74ffc9ea
Patrick von Platen authored May 07, 2020
```
* fix example reformer

* fix error message and example docstring

* improved error message
```
74ffc9ea
fix docstring reformer (#4190) · 96c78396
Patrick von Platen authored May 07, 2020

96c78396

Reformer (#3351) · dca34695

Patrick von Platen authored May 07, 2020

* first copy & past commit from Bert and morgans LSH code

* add easy way to compare to trax original code

* translate most of function

* make trax lsh self attention deterministic with numpy seed + copy paste code

* add same config

* add same config

* make layer init work

* implemented hash_vectors function for lsh attention

* continue reformer translation

* hf LSHSelfAttentionLayer gives same output as trax layer

* refactor code

* refactor code

* refactor code

* refactor

* refactor + add reformer config

* delete bogus file

* split reformer attention layer into two layers

* save intermediate step

* save intermediate step

* make test work

* add complete reformer block layer

* finish reformer layer

* implement causal and self mask

* clean reformer test and refactor code

* fix merge conflicts

* fix merge conflicts

* update init

* fix device for GPU

* fix chunk length init for tests

* include morgans optimization

* improve memory a bit

* improve comment

* factorize num_buckets

* better testing parameters

* make whole model work

* make lm model work

* add t5 copy paste tokenizer

* add chunking feed forward

* clean config

* add improved assert statements

* make tokenizer work

* improve test

* correct typo

* extend config

* add complexer test

* add new axial position embeddings

* add local block attention layer

* clean tests

* refactor

* better testing

* save intermediate progress

* clean test file

* make shorter input length work for model

* allow variable input length

* refactor

* make forward pass for pretrained model work

* add generation possibility

* finish dropout and init

* make style

* refactor

* add first version of RevNet Layers

* make forward pass work and add convert file

* make uploaded model forward pass work

* make uploaded model forward pass work

* refactor code

* add namedtuples and cache buckets

* correct head masks

* refactor

* made reformer more flexible

* make style

* remove set max length

* add attention masks

* fix up tests

* fix lsh attention mask

* make random seed optional for the moment

* improve memory in reformer

* add tests

* make style

* make sure masks work correctly

* detach gradients

* save intermediate

* correct backprob through gather

* make style

* change back num hashes

* rename to labels

* fix rotation shape

* fix detach

* update

* fix trainer

* fix backward dropout

* make reformer more flexible

* fix conflict

* fix

* fix

* add tests for fixed seed in reformer layer

* fix trainer typo

* fix typo in activations

* add fp16 tests

* add fp16 training

* support fp16

* correct gradient bug in reformer

* add fast gelu

* re-add dropout for embedding dropout

* better naming

* better naming

* renaming

* finalize test branch

* finalize tests

* add more tests

* finish tests

* fix

* fix type trainer

* fix fp16 tests

* fix tests

* fix tests

* fix tests

* fix issue with dropout

* fix dropout seeds

* correct random seed on gpu

* finalize random seed for dropout

* finalize random seed for dropout

* remove duplicate line

* correct half precision bug

* make style

* refactor

* refactor

* docstring

* remove sinusoidal position encodings for reformer

* move chunking to modeling_utils

* make style

* clean config

* make style

* fix tests

* fix auto tests

* pretrained models

* fix docstring

* update conversion file

* Update pretrained_models.rst

* fix rst

* fix rst

* update copyright

* fix test path

* fix test path

* fix small issue in test

* include reformer in generation tests

* add docs for axial position encoding

* finish docs

* Update convert_reformer_trax_checkpoint_to_pytorch.py

* remove isort

* include sams comments

* remove wrong comment in utils

* correct typos

* fix typo

* Update reformer.rst

* applied morgans optimization

* make style

* make gpu compatible

* remove bogus file

* big test refactor

* add example for chunking

* fix typo

* add to README

dca34695

06 May, 2020 7 commits

change order pytorch/tf in readme (#4167) · 877fc564
Clement authored May 06, 2020

877fc564

TF version of the trainer (#4017) · aad50151

Julien Plu authored May 06, 2020

* First commit to add a TF version of the trainer.

* Make the TF trainer closer to what looks the PT trainer

* Refactoring common code between the PT and TF trainer into an util file.

* Some bugfix + better similarity with the PT trainer

* Add missing class in transformers init

* Bugfix over prediction + use classification report instead of simple metrics

* Fix name error

* Fix optimization tests + style

* Apply style

* Several bugfix for multi-gpu training

* Apply style

* Apply style

* Add glue example for the TF trainer

* Several bugix + address the reviews

* Fix on the TF training args file

* Add a debug mode

* Bugfix in utils_ner.py when segment_ids is None

* Apply style

* Apply style

* Add TPU strategy

* Fix selection strategy

aad50151

Fix overwrite_cache behaviour for pytorch lightning examples (#4093) · 25296b12
Simone Primarosa authored May 06, 2020

25296b12
Include ElectraPreTrainedModel into __init__ (#4173) · 9972562d
kumapo authored May 07, 2020

9972562d
Camembert-large-fquad model card (#4143) · ff8ed52d
martindh authored May 06, 2020
```
Description for the model card describing the camembert-large-fquad model.
```
ff8ed52d
Add model card for the NER model (#4162) · 4c3be2e7
Julien Plu authored May 06, 2020

4c3be2e7
Fix markdown to show the results table properly (#4119) · 17ae0363
Manuel Romero authored May 06, 2020

17ae0363

05 May, 2020 4 commits

fix hard wired pad token id (#4138) · a638e986
Patrick von Platen authored May 06, 2020

a638e986
[Trainer] W&B: Enable model watch · fd217466
Julien Chaumond authored May 05, 2020
```
See https://github.com/huggingface/transformers/pull/3916
```
fd217466

Pytorch 1.5.0 (#3973) · 79b1c696

Lysandre Debut authored May 05, 2020

* Standard deviation can no longer be set to 0

* Remove torch pinned version

* 9th instead of 10th, silly me

79b1c696

Trainer: add logging through Weights & Biases (#3916) · 818463ee

Boris Dayma authored May 04, 2020



* feat: add logging through Weights & Biases

* feat(wandb): make logging compatible with all scripts

* style(trainer.py): fix formatting

* [Trainer] Tweak wandb integration
Co-authored-by: Julien Chaumond <chaumond@gmail.com>

818463ee

04 May, 2020 2 commits

allow an already created tensorboard SummaryWriter be passed to Trainer · 858b1d1e
jaymody authored Apr 27, 2020

858b1d1e

[EncoderDecoder Tests] Improve tests (#4046) · 8e67573a

Patrick von Platen authored May 04, 2020



* Hoist bert model tester for patric

* indent

* make tests work

* Update tests/test_modeling_bert.py
Co-authored-by: Julien Chaumond <chaumond@gmail.com>
Co-authored-by: sshleifer <sshleifer@gmail.com>
Co-authored-by: Julien Chaumond <chaumond@gmail.com>

8e67573a

03 May, 2020 1 commit
- Add decoder specific error message for T5Stack.forward (#4128) · 6af3306a
  Lorenzo Ampil authored May 03, 2020
  
  6af3306a
02 May, 2020 9 commits
- Fix #2941 (#4109) · 1cdd2ad2
  Zhiyu Lin authored May 02, 2020
```
* Fix of issue #2941

Reshaped score array to avoid `numpy` ValueError.

* Update src/transformers/pipelines.py

* Update src/transformers/pipelines.py
Co-authored-by: Julien Chaumond <chaumond@gmail.com>
```
  1cdd2ad2
- distilroberta-base-finetuned-sentiment (#4115) · 5f4f6b65
  Manuel Romero authored May 02, 2020
```
* Create model card

Create Model card for distilroberta-base-finetuned-sentiment

* Update model_cards/mrm8488/distilroberta-base-finetuned-sentiment/README.md

* Update model_cards/mrm8488/distilroberta-base-finetuned-sentiment/README.md
Co-authored-by: Julien Chaumond <chaumond@gmail.com>
```
  5f4f6b65
- model card for surajp/albert-base-sanskrit (#4114) · 7da051f1
  Suraj Parmar authored May 02, 2020
```
* Create README.md

* Update model_cards/surajp/albert-base-sanskrit/README.md
Co-authored-by: Julien Chaumond <chaumond@gmail.com>
```
  7da051f1
- Create README.md (#4112) · 14911e2e
  Zhen Wang authored May 02, 2020
  
  14911e2e
- Added huseinzol05/gpt2-345M-bahasa-cased (#4102) · 9e97c875
  HUSEIN ZOLKEPLI authored May 02, 2020
  
  9e97c875
- Update run_pl_glue.py (#4117) · 4c5bd921
  William Falcon authored May 02, 2020
  
  4c5bd921
- Update run_pl_ner.py (#4118) · 5282b31d
  William Falcon authored May 02, 2020
  
  5282b31d
- NER: parse args from .args file or JSON (#4110) · 1e616c0a
  Stefan Schweter authored May 02, 2020
```
* ner: parse args from .args file or JSON

* examples: mention json-based configuration file support for run_ner script
```
  1e616c0a
- Update README.md · abb1fa3f
  Patrick von Platen authored May 02, 2020
  
  abb1fa3f