Commits · c8035e11e88731bea98b092c18b5bff6c7589c18 · chenpangpang / transformers

05 Mar, 2020 8 commits
- fix renaming problem · 58fc8f97
  patrickvonplaten authored Mar 06, 2020
  
  58fc8f97
- Rename BartForMaskedLM -> BartForConditionalGeneration (#3114) · 857e0a0d
  Sam Shleifer authored Mar 05, 2020
```
* improved documentation
```
  857e0a0d
- Pass kwargs to configuration (#3147) · b623ddc0
  Lysandre Debut authored Mar 05, 2020
```
* Pass kwargs to configuration

* Setter

* test
```
  b623ddc0
- Correct missing keys + test (#3143) · 0001d056
  Lysandre Debut authored Mar 05, 2020
  
  0001d056
- cleanup deltas · 1360daca
  sshleifer authored Mar 05, 2020
  
  1360daca
- tests pass · c36fdc88
  sshleifer authored Mar 05, 2020
  
  c36fdc88
- make style · ff9e79ba
  Julien Chaumond authored Mar 04, 2020
  
  ff9e79ba
- Fix failing doc samples · 07a79db5
  Lysandre authored Mar 04, 2020
  
  07a79db5
04 Mar, 2020 2 commits
- include tf gpt2 tests for attn mask and past variable (#3122) · 932eab94
  Patrick von Platen authored Mar 04, 2020
  
  932eab94
- added beam_search generation for tf 2.0 · 61fef6e9
  patrickvonplaten authored Mar 04, 2020
  
  61fef6e9
03 Mar, 2020 3 commits

BartForSequenceClassification: fix num_labels, add test (#3110) · e9e6efdc
Sam Shleifer authored Mar 03, 2020

e9e6efdc

[ci] Re-run integration ground truth from fairseq · f631e01d

Julien Chaumond authored Mar 03, 2020

Adopted best practice set by @patrickvonplaten of commenting lines run on fairseq, for easy comparison

also see #3020

f631e01d

Add generate() functionality to TF 2.0 (#3063) · 41341003

Patrick von Platen authored Mar 03, 2020

* add first copy past test to tf 2 generate

* add tf top_k_top_p_filter fn

* add generate function for TF

* add generate function for TF

* implemented generate for all models expect transfoXL

* implemented generate for all models expect transfoXL

* implemented generate for all models expect transfoXL

* make style

* change permission of test file to correct ones

* delete ipdb

* delete ipdb

* fix bug and finish simple gpt2 integration test

* clean test file

* clean test file

* make style

* make style

* make style

* make style

* change import style

* change import style

* make style

* make style

* add decorators

* add decorators

* fix tf ctrl bug dim => axis in TF

* make style

* make style

* refactored test file

* refactored test file

* take out test_torch_tf_conversion if nothing is defined

* take out test_torch_tf_conversion if nothing is defined

* remove useless files

* remove useless files

* fix conflicts

* fix conflicts

* fix conflicts

* fix conflicts

* fix conflicts

* solve conflicts

* solve conflicts

* fix conflicts

* fix conflicts

* merge conflicts

* delete ipdb

* exposed top_k_top_p_filtering fns

* delete weirdly created w! file

* add comment to test tf common modeling

* fix conflicts

* fix conflicts

* make style

* merge conflicts

* make style

* change tf.tensor.shape to shape_list(tensor)

41341003

02 Mar, 2020 6 commits

TF GPU CI (#3085) · f169957d

Julien Chaumond authored Mar 02, 2020

* debug env

* Restrict TF GPU memory

* Fixup

* One more test

* rm debug logs

* Fixup

f169957d

Pipeline doc (#3055) · d3eb7d23

Lysandre Debut authored Mar 02, 2020

* Pipeline doc initial commit

* pipeline abstraction

* Remove modelcard argument from pipeline

* Task-specific pipelines can be instantiated with no model or tokenizer

* All pipelines doc

d3eb7d23

rm bogus file · 0e56b37e
Julien Chaumond authored Mar 02, 2020
```
cc @patrickvonplaten
```
0e56b37e
correct greedy generation when doing beam search (#3078) · 2fdc7f6c
Patrick von Platen authored Mar 02, 2020
```
* correct greedy generation when doing beam search

* improve comment
```
2fdc7f6c
Force pad_token_id to be set before padding for standard tokenizer (#3035) · c0135194
Patrick von Platen authored Mar 02, 2020
```
* force pad_token_id to be set before padding

* fix tests and forbid padding without having a padding_token_id set
```
c0135194

Bart-CNN (#3059) · b54ef78d

Sam Shleifer authored Mar 02, 2020

`generate` code that produces 99% identical summarizations to fairseq on CNN test data, with caching.

b54ef78d

27 Feb, 2020 2 commits
- Added test for AlbertForTokenClassification · f7115752
  Martin Malmsten authored Feb 27, 2020
  
  f7115752
- Added test for AlbertForTokenClassification · aceb6a09
  Martin Malmsten authored Feb 27, 2020
  
  aceb6a09
26 Feb, 2020 5 commits
- [gpu] Fixup fdd61b19 · b370cc7e
  Julien Chaumond authored Feb 26, 2020
  
  b370cc7e
- Fix bart slow test · f5516805
  Julien Chaumond authored Feb 26, 2020
  
  f5516805
- Fix attn mask gpt2 when using past (#3033) · fdd61b19
  Patrick von Platen authored Feb 26, 2020
```
* fix issue and add some tests

* fix issue and add some tests

* updated doc string gpt2
```
  fdd61b19
- Fix (non-slow) tests on GPU (torch) (#3024) · 9cda3620
  Julien Chaumond authored Feb 26, 2020
```
* Fix tests on GPU (torch)

* Fix bart slow tests
Co-authored-by: Sam Shleifer <sshleifer@gmail.com>
```
  9cda3620
- Delete all mentions of Model2Model (#3019) · 9df74b8b
  Sam Shleifer authored Feb 26, 2020
  
  9df74b8b
25 Feb, 2020 3 commits
- Add integration tests for xlm roberta modelling and xlm roberta tokenzier (#3014) · c913eb9c
  Patrick von Platen authored Feb 25, 2020
```
* add first files

* add xlm roberta integration tests

* make style

* flake 8 issues solved
```
  c913eb9c
- make style · f5b50c6b
  Patrick von Platen authored Feb 25, 2020
  
  f5b50c6b
- add special tokens to pretrain configs of respective lm head models · e645dcbb
  Patrick von Platen authored Feb 25, 2020
  
  e645dcbb
24 Feb, 2020 4 commits

Test correct tokenizers after default switch (#3003) · b90745c5
Lysandre Debut authored Feb 24, 2020

b90745c5

Fix for fast tokenizers save_pretrained compatibility with Python. (#2933) · 4cd9c097

Funtowicz Morgan authored Feb 25, 2020



* Renamed file generate by tokenizers when calling save_pretrained to match python.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Added save_vocabulary tests.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Remove python quick and dirty fix for clean Rust impl.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Bump tokenizers dependency to 0.5.1
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* TransfoXLTokenizerFast uses a json vocabulary file + warning about incompatibility between Python and Rust
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Added some save_pretrained / from_pretrained unittests.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Update tokenizers to 0.5.2
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Quality and format.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* flake8
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Making sure there is really a bug in unittest

* Fix TransfoXL constructor vocab_file / pretrained_vocab_file mixin.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

4cd9c097

Testing that batch_encode_plus is the same as encode_plus (#2973) · 21d8b6a3

Lysandre Debut authored Feb 24, 2020

* Testing that encode_plus and batch_encode_plus behave the same way

Spoiler alert: they don't

* Testing rest of arguments in batch_encode_plus

* Test tensor return in batch_encode_plus

* Addressing Sam's comments

* flake8

* Simplified with `num_added_tokens`

21d8b6a3

Add slow generate tests for pretrained lm models (#2909) · 17c45c39

Patrick von Platen authored Feb 24, 2020

* add slow generate lm_model tests

* fix conflicts

* merge conflicts

* fix conflicts

* add slow generate lm_model tests

* make style

* delete unused variable

* fix conflicts

* fix conflicts

* fix conflicts

* delete unused variable

* fix conflicts

* finished hard coded tests

17c45c39

22 Feb, 2020 2 commits
- Bart: fix layerdrop and cached decoder_input_ids for generation (#2969) · 92487a1d
  Sam Shleifer authored Feb 22, 2020
  
  92487a1d
- Fix max_length not taken into account when using pad_to_max_length on fast tokenizers (#2961) · cc6775cd
  Funtowicz Morgan authored Feb 22, 2020
```
* enable_padding should pad up to max_length if set.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Added more testing on padding.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>
```
  cc6775cd
21 Feb, 2020 1 commit

Improve special_token_id logic in run_generation.py and add tests (#2885) · fc38d4c8

Patrick von Platen authored Feb 21, 2020



* improving generation

* finalized special token behaviour for no_beam_search generation

* solved modeling_utils merge conflict

* solve merge conflicts in modeling_utils.py

* add run_generation improvements from PR #2749

* adapted language generation to not use hardcoded -1 if no padding token is available

* remove the -1 removal as hard coded -1`s are not necessary anymore

* add lightweight language generation testing for randomely initialized models - just checking whether no errors are thrown

* add slow language generation tests for pretrained models using hardcoded output with pytorch seed

* delete ipdb

* check that all generated tokens are valid

* renaming

* renaming Generation -> Generate

* make style

* updated so that generate_beam_search has same token behavior than generate_no_beam_search

* consistent return format for run_generation.py

* deleted pretrain lm generate tests -> will be added in another PR

* cleaning of unused if statements and renaming

* run_generate will always return an iterable

* make style

* consistent renaming

* improve naming, make sure generate function always returns the same tensor, add docstring

* add slow tests for all lmhead models

* make style and improve example comments modeling_utils

* better naming and refactoring in modeling_utils

* improving generation

* finalized special token behaviour for no_beam_search generation

* solved modeling_utils merge conflict

* solve merge conflicts in modeling_utils.py

* add run_generation improvements from PR #2749

* adapted language generation to not use hardcoded -1 if no padding token is available

* remove the -1 removal as hard coded -1`s are not necessary anymore

* add lightweight language generation testing for randomely initialized models - just checking whether no errors are thrown

* add slow language generation tests for pretrained models using hardcoded output with pytorch seed

* delete ipdb

* check that all generated tokens are valid

* renaming

* renaming Generation -> Generate

* make style

* updated so that generate_beam_search has same token behavior than generate_no_beam_search

* consistent return format for run_generation.py

* deleted pretrain lm generate tests -> will be added in another PR

* cleaning of unused if statements and renaming

* run_generate will always return an iterable

* make style

* consistent renaming

* improve naming, make sure generate function always returns the same tensor, add docstring

* add slow tests for all lmhead models

* make style and improve example comments modeling_utils

* better naming and refactoring in modeling_utils

* changed fast random lm generation testing design to more general one

* delete in old testing design in gpt2

* correct old variable name

* temporary fix for encoder_decoder lm generation tests - has to be updated when t5 is fixed

* adapted all fast random generate tests to new design

* better warning description in modeling_utils

* better comment

* better comment and error message
Co-authored-by: Thomas Wolf <thomwolf@users.noreply.github.com>

fc38d4c8

20 Feb, 2020 2 commits
- New BartModel (#2745) · 53ce3854
  Sam Shleifer authored Feb 20, 2020
```
* Results same as fairseq
* Wrote a ton of tests
* Struggled with api signatures
* added some docs
```
  53ce3854
- Add get_vocab method to PretrainedTokenizer · 197d74f9
  Joe Davison authored Feb 20, 2020
  
  197d74f9
19 Feb, 2020 2 commits

Fast Tokenizers save pretrained should return the list of generated file paths. (#2918) · d490b5d5

Funtowicz Morgan authored Feb 20, 2020



* Correctly return the tuple of generated file(s) when calling save_pretrained
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Quality and format.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

d490b5d5

Override build_inputs_with_special_tokens for fast tokenizers (#2912) · e6767642

Funtowicz Morgan authored Feb 19, 2020



* Override build_inputs_with_special_tokens for fast impl + unittest.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Quality + format.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

e6767642