Commits · fc38d4c86fe4bbde91b194880fe38b821a346123 · chenpangpang / transformers

21 Feb, 2020 5 commits

Improve special_token_id logic in run_generation.py and add tests (#2885) · fc38d4c8

Patrick von Platen authored Feb 21, 2020



* improving generation

* finalized special token behaviour for no_beam_search generation

* solved modeling_utils merge conflict

* solve merge conflicts in modeling_utils.py

* add run_generation improvements from PR #2749

* adapted language generation to not use hardcoded -1 if no padding token is available

* remove the -1 removal as hard coded -1`s are not necessary anymore

* add lightweight language generation testing for randomely initialized models - just checking whether no errors are thrown

* add slow language generation tests for pretrained models using hardcoded output with pytorch seed

* delete ipdb

* check that all generated tokens are valid

* renaming

* renaming Generation -> Generate

* make style

* updated so that generate_beam_search has same token behavior than generate_no_beam_search

* consistent return format for run_generation.py

* deleted pretrain lm generate tests -> will be added in another PR

* cleaning of unused if statements and renaming

* run_generate will always return an iterable

* make style

* consistent renaming

* improve naming, make sure generate function always returns the same tensor, add docstring

* add slow tests for all lmhead models

* make style and improve example comments modeling_utils

* better naming and refactoring in modeling_utils

* improving generation

* finalized special token behaviour for no_beam_search generation

* solved modeling_utils merge conflict

* solve merge conflicts in modeling_utils.py

* add run_generation improvements from PR #2749

* adapted language generation to not use hardcoded -1 if no padding token is available

* remove the -1 removal as hard coded -1`s are not necessary anymore

* add lightweight language generation testing for randomely initialized models - just checking whether no errors are thrown

* add slow language generation tests for pretrained models using hardcoded output with pytorch seed

* delete ipdb

* check that all generated tokens are valid

* renaming

* renaming Generation -> Generate

* make style

* updated so that generate_beam_search has same token behavior than generate_no_beam_search

* consistent return format for run_generation.py

* deleted pretrain lm generate tests -> will be added in another PR

* cleaning of unused if statements and renaming

* run_generate will always return an iterable

* make style

* consistent renaming

* improve naming, make sure generate function always returns the same tensor, add docstring

* add slow tests for all lmhead models

* make style and improve example comments modeling_utils

* better naming and refactoring in modeling_utils

* changed fast random lm generation testing design to more general one

* delete in old testing design in gpt2

* correct old variable name

* temporary fix for encoder_decoder lm generation tests - has to be updated when t5 is fixed

* adapted all fast random generate tests to new design

* better warning description in modeling_utils

* better comment

* better comment and error message
Co-authored-by: Thomas Wolf <thomwolf@users.noreply.github.com>

fc38d4c8

Added CamembertForQuestionAnswering (#2746) · c749a543
maximeilluin authored Feb 21, 2020
```
* Added CamembertForQuestionAnswering

* fixed camembert tokenizer case
```
c749a543

Update modeling_tf_utils.py (#2924) · 5211d333

Bram Vanroy authored Feb 21, 2020

Tensorflow does not use .eval() vs .train().

closes https://github.com/huggingface/transformers/issues/2906

5211d333

Create README.md for xlnet_large_squad (#2942) · 3e98f27e
ahotrod authored Feb 21, 2020

3e98f27e
Labels are now added to model config under id2label and label2id (#2945) · 4452b44b
Martin Malmsten authored Feb 21, 2020

4452b44b

20 Feb, 2020 10 commits
- New BartModel (#2745) · 53ce3854
  Sam Shleifer authored Feb 20, 2020
```
* Results same as fairseq
* Wrote a ton of tests
* Struggled with api signatures
* added some docs
```
  53ce3854
- Removed unused fields in DistilBert TransformerBlock (#2710) · 564fd75d
  guillaume-be authored Feb 20, 2020
```
* Removed unused fields in DistilBert TransformerBlock
```
  564fd75d
- default arg fix (#2937) · 889d3bfd
  srush authored Feb 20, 2020
  
  889d3bfd
- Fix InputExample docstring (#2891) · ea8eba35
  Scott Gigante authored Feb 20, 2020
  
  ea8eba35
- Tokenizer fast warnings (#2922) · e2a6445e
  Funtowicz Morgan authored Feb 20, 2020
```
* Remove warning when pad_to_max_length is not set.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Move RoberTa warning to RoberTa and not GPT2 base tokenizer.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>
```
  e2a6445e
- Expose all constructor parameter for BertTokenizerFast (#2921) · 9b309331
  Funtowicz Morgan authored Feb 20, 2020
```
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>
```
  9b309331
- Support for torch-lightning in NER examples (#2890) · b662f0e6
  srush authored Feb 20, 2020
```
* initial pytorch lightning commit

* tested multigpu

* Fix learning rate schedule

* black formatting

* fix flake8

* isort

* isort

* .
Co-authored-by: Check your git settings! <chris@chris-laptop>
```
  b662f0e6
- Update to include example of LM · ab123839
  Ilias Chalkidis authored Feb 19, 2020
```
The model files have been updated in order to include the classification layers, based on https://github.com/huggingface/transformers/issues/2901, and now can be also used as a LM.
```
  ab123839
- Add syntax highlighting to the BibTeX in README · 976e9afe
  Santiago Castro authored Feb 19, 2020
  
  976e9afe
- Fix spell: EsperBERTo, not EspertBERTo · cbc57055
  Cong authored Feb 20, 2020
  
  cbc57055
19 Feb, 2020 9 commits

Fast Tokenizers save pretrained should return the list of generated file paths. (#2918) · d490b5d5

Funtowicz Morgan authored Feb 20, 2020



* Correctly return the tuple of generated file(s) when calling save_pretrained
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Quality and format.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

d490b5d5

Patch ALBERT with heads in TensorFlow · 2708b44e
Lysandre authored Feb 19, 2020

2708b44e
Patch ALBERT with heads in TensorFlow · 1abd53b1
Lysandre authored Feb 19, 2020

1abd53b1

Override build_inputs_with_special_tokens for fast tokenizers (#2912) · e6767642

Funtowicz Morgan authored Feb 19, 2020



* Override build_inputs_with_special_tokens for fast impl + unittest.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Quality + format.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

e6767642

README link + better instructions for release · 59c23ad9
Lysandre authored Feb 19, 2020

59c23ad9
Documentation v2.5.0 · 22b2b579
Lysandre authored Feb 19, 2020

22b2b579
Release: v2.5.0 · fb560dcb
Lysandre authored Feb 19, 2020
```
Welcome Rust Tokenizers
```
fb560dcb

Integrate fast tokenizers library inside transformers (#2674) · 3f3fa7f7

Funtowicz Morgan authored Feb 19, 2020



* Implemented fast version of tokenizers
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Bumped tokenizers version requirements to latest 0.2.1
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Added matching tests
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Matching OpenAI GPT tokenization !
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Matching GPT2 on tokenizers
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Expose add_prefix_space as constructor parameter for GPT2
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Matching Roberta tokenization !
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Removed fast implementation of CTRL.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Binding TransformerXL tokenizers to Rust.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Updating tests accordingly.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Added tokenizers as top-level modules.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Black & isort.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Rename LookupTable to WordLevel to match Rust side.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Black.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Use "fast" suffix instead of "ru" for rust tokenizers implementations.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Introduce tokenize() method on fast tokenizers.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* encode_plus dispatchs to batch_encode_plus
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* batch_encode_plus now dispatchs to encode if there is only one input element.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Bind all the encode_plus parameter to the forwarded batch_encode_plus call.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Bump tokenizers dependency to 0.3.0
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Formatting.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Fix tokenization_auto with support for new (python, fast) mapping schema.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Give correct fixtures path in test_tokenization_fast.py for the CLI.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Expose max_len_ properties on BertTokenizerFast
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Move max_len_ properties to PreTrainedTokenizerFast and override in specific subclasses.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* _convert_encoding should keep the batch axis tensor if only one sample in the batch.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Add warning message for RobertaTokenizerFast if used for MLM.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Added use_fast (bool) parameter on AutoTokenizer.from_pretrained().

This allows to easily enable/disable Rust-based tokenizer instantiation.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Let's tokenizers handle all the truncation and padding stuff.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Allow to provide tokenizer arguments during pipeline creation.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Update test_fill_mask pipeline to not use fast tokenizers.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Fix too much parameters for convert_encoding.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* When enabling padding, max_length should be set to None.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Avoid returning nested tensors of length 1 when calling encode_plus
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Ensure output is padded when return_tensor is not None.

Tensor creation requires the inital list input to be of the exact same size.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Disable transfoxl unittest if pytorch is not available (required to load the model)
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* encode_plus should not remove the leading batch axis if return_tensor is set
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Temporary disable fast tokenizers on QA pipelines.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Fix formatting issues.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Update tokenizers to 0.4.0

* Update style

* Enable truncation + stride unit test on fast tokenizers.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Add unittest ensuring special_tokens set match between Python and Rust.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Ensure special_tokens are correctly set during construction.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Give more warning feedback to the user in case of padding without pad_token.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* quality & format.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Added possibility to add a single token as str
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Added unittest for add_tokens and add_special_tokens on fast tokenizers.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Fix rebase mismatch on pipelines qa default model.

QA requires cased input while the tokenizers would be uncased.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Addressing review comment: Using offset mapping relative to the original string + unittest.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Addressing review comment: save_vocabulary requires folder and file name
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Addressing review comment: Simplify import for Bert.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Addressing review comment: truncate_and_pad disables padding according to the same heuristic than the one enabling padding.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Addressing review comment: Remove private member access in tokenize()
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Addressing review comment: Bump tokenizers dependency to 0.4.2
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* format & quality.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Addressing review comment: Use named arguments when applicable.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Addressing review comment: Add Github link to Roberta/GPT2 space issue on masked input.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Addressing review comment: Move max_len_single_sentence / max_len_sentences_pair to PreTrainedTokenizerFast + tests.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Addressing review comment: Relax type checking to include tuple and list object.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Addressing review comment: Document the truncate_and_pad manager behavior.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Raise an exception if return_offsets_mapping is not available with the current tokenizer.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Ensure padding is set on the tokenizers before setting any padding strategy + unittest.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* On pytorch we need to stack tensor to get proper new axis.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Generalize tests to different framework removing hard written return_tensors="..."
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Bump tokenizer dependency for num_special_tokens_to_add
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Overflowing tokens in batch_encode_plus are now stacked over the batch axis.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Improved error message for padding strategy without pad token.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Bumping tokenizers dependency to 0.5.0 for release.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Optimizing convert_encoding around 4x improvement. 🚀

Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* expose pad_to_max_length in encode_plus to avoid duplicating the parameters in kwargs
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Generate a proper overflow_to_sampling_mapping when return_overflowing_tokens is True.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Fix unittests for overflow_to_sampling_mapping not being returned as tensor.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Format & quality.
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Remove perfect alignment constraint for Roberta (allowing 1% difference max)
Signed-off-by: Morgan Funtowicz <morgan@huggingface.co>

* Triggering final CI
Co-authored-by: MOI Anthony <xn1t0x@gmail.com>

3f3fa7f7

Create README.md · ffb93ec0
Bin Wang authored Feb 18, 2020

ffb93ec0

18 Feb, 2020 2 commits
- Skip flaky test_tf_question_answering (#2845) · 20fc18fb
  Sam Shleifer authored Feb 18, 2020
```
* Skip flaky test

* Style
```
  20fc18fb
- fix vocab size in binarized_data (distil): int16 vs int32 · 2ae98336
  VictorSanh authored Feb 18, 2020
  
  2ae98336
17 Feb, 2020 5 commits
- fix typo in hans example call · 0dbddba6
  VictorSanh authored Feb 17, 2020
  
  0dbddba6
- Create README.md · 29ab4b7f
  Manuel Romero authored Feb 17, 2020
  
  29ab4b7f
- [model_cards] 🇹🇷 Add new (cased) BERTurk model · c88ed74c
  Stefan Schweter authored Feb 17, 2020
  
  c88ed74c
- Merge pull request #2881 from patrickvonplaten/add_vim_swp_to_gitignore · 5b2d4f26
  Thomas Wolf authored Feb 17, 2020
```
update .gitignore to ignore .swp files created when using vim
```
  5b2d4f26
- update .gitignore to ignore .swp files created when using vim · fb4d8d08
  Patrick von Platen authored Feb 17, 2020
  
  fb4d8d08
16 Feb, 2020 1 commit

Update README.md · 6083c156

Manuel Romero authored Feb 16, 2020

I trained the model for more epochs so I improved the results. This commit will update the results of the model and add a gif using it with **transformers/pipelines**

6083c156

14 Feb, 2020 8 commits
- [model_cards] EsperBERTo · 73028c5d
  Julien Chaumond authored Feb 14, 2020
  
  73028c5d
- Update model card: new performance chart (#2864) · 81fb8d32
  Timo Moeller authored Feb 14, 2020
```
* Update model performance for correct German conll03 dataset

* Adjust text

* Adjust line spacing
```
  81fb8d32
- [model_cards] Also use the thumbnail as meta · 4e69104a
  Julien Chaumond authored Feb 14, 2020
```
Co-Authored-By: Ilias Chalkidis <ihalk@di.uoa.gr>
```
  4e69104a
- [model_cards] nlptown/bert-base-multilingual-uncased-sentiment · 73d79d42
  Julien Chaumond authored Feb 14, 2020
```
cc @yvespeirsman
Co-Authored-By: Yves Peirsman <yvespeirsman@users.noreply.github.com>
```
  73d79d42
- Added model card for bert-base-multilingual-uncased-sentiment (#2859) · 47b735f9
  Yves Peirsman authored Feb 14, 2020
```
* Created model card for nlptown/bert-base-multilingual-sentiment

* Delete model card

* Created model card for bert-base-multilingual-uncased-sentiment as README
```
  47b735f9
- [pipeline] Alias NerPipeline as TokenClassificationPipeline · 7d22fefd
  Julien Chaumond authored Feb 13, 2020
  
  7d22fefd
- Fix typo · 61a2b7dc
  Manuel Romero authored Feb 14, 2020
  
  61a2b7dc
- Fix typos · 6e261d3a
  Ilias Chalkidis authored Feb 14, 2020
  
  6e261d3a