Commits · ba8c4d0ac04acfcdbdeaed954f698d6d5ec3e532 · chenpangpang / transformers

18 Oct, 2020 1 commit

[Dependencies|tokenizers] Make both SentencePiece and Tokenizers optional dependencies (#7659) · ba8c4d0a

Thomas Wolf authored Oct 18, 2020

* splitting fast and slow tokenizers [WIP]

* [WIP] splitting sentencepiece and tokenizers dependencies

* update dummy objects

* add name_or_path to models and tokenizers

* prefix added to file names

* prefix

* styling + quality

* spliting all the tokenizer files - sorting sentencepiece based ones

* update tokenizer version up to 0.9.0

* remove hard dependency on sentencepiece 🎉

* and removed hard dependency on tokenizers 🎉



* update conversion script

* update missing models

* fixing tests

* move test_tokenization_fast to main tokenization tests - fix bugs

* bump up tokenizers

* fix bert_generation

* update ad fix several tokenizers

* keep sentencepiece in deps for now

* fix funnel and deberta tests

* fix fsmt

* fix marian tests

* fix layoutlm

* fix squeezebert and gpt2

* fix T5 tokenization

* fix xlnet tests

* style

* fix mbart

* bump up tokenizers to 0.9.2

* fix model tests

* fix tf models

* fix seq2seq examples

* fix tests without sentencepiece

* fix slow => fast  conversion without sentencepiece

* update auto and bert generation tests

* fix mbart tests

* fix auto and common test without tokenizers

* fix tests without tokenizers

* clean up tests lighten up when tokenizers + sentencepiece are both off

* style quality and tests fixing

* add sentencepiece to doc/examples reqs

* leave sentencepiece on for now

* style quality split hebert and fix pegasus

* WIP Herbert fast

* add sample_text_no_unicode and fix hebert tokenization

* skip FSMT example test for now

* fix style

* fix fsmt in example tests

* update following Lysandre and Sylvain's comments

* Update src/transformers/testing_utils.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/testing_utils.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/tokenization_utils_base.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Update src/transformers/tokenization_utils_base.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

ba8c4d0a

31 Aug, 2020 1 commit
- Fix marian slow test (#6854) · 8af1970e
  Sam Shleifer authored Aug 31, 2020
  
  8af1970e
24 Aug, 2020 1 commit
- Update repo to isort v5 (#6686) · a5737779
  Sylvain Gugger authored Aug 24, 2020
```
* Run new isort

* More changes

* Update CI, CONTRIBUTING and benchmarks
```
  a5737779
19 Aug, 2020 1 commit
- Fix bart base test (#6587) · ab42d748
  Sam Shleifer authored Aug 18, 2020
  
  ab42d748
18 Aug, 2020 1 commit
- [marian] converter supports models from new Tatoeba project (#6342) · 12d76241
  Sam Shleifer authored Aug 17, 2020
  
  12d76241
11 Aug, 2020 1 commit
- rename prepare_translation_batch -> prepare_seq2seq_batch (#6103) · be1520d3
  Sam Shleifer authored Aug 11, 2020
  
  be1520d3
01 Jul, 2020 1 commit
- Move tests/utils.py -> transformers/testing_utils.py (#5350) · 13deb95a
  Sam Shleifer authored Jul 01, 2020
  
  13deb95a
06 Jun, 2020 1 commit
- [marian tests ] pass device to pipeline (#4815) · c58e6c12
  Sam Shleifer authored Jun 06, 2020
  
  c58e6c12
05 Jun, 2020 1 commit
- [cleanup/marian] pipelines test and new kwarg (#4812) · 4ab74245
  Sam Shleifer authored Jun 05, 2020
  
  4ab74245
19 May, 2020 1 commit
- [MarianTokenizer] implement save_vocabulary and other common methods (#4389) · efbc1c5a
  Sam Shleifer authored May 19, 2020
  
  efbc1c5a
13 May, 2020 1 commit
- [Marian Fixes] prevent predicting pad_token_id before softmax, support... · 9a687ebb
  Sam Shleifer authored May 13, 2020
```
[Marian Fixes] prevent predicting pad_token_id before softmax, support language codes, name multilingual models (#4290)
```
  9a687ebb
10 May, 2020 1 commit

[Marian] documentation and AutoModel support (#4152) · 3487be75

Sam Shleifer authored May 10, 2020

- MarianSentencepieceTokenizer - > MarianTokenizer
- Start using unk token.
- add docs page
- add better generation params to MarianConfig
- more conversion utilities

3487be75

28 Apr, 2020 1 commit
- MarianMTModel.from_pretrained('Helsinki-NLP/opus-marian-en-de') (#3908) · 847e7f33
  Sam Shleifer authored Apr 28, 2020
```
Co-Authored-By: Stefan Schweter <stefan@schweter.it>
```
  847e7f33